Background
Five anonymous programmers brought a proposed class action against GitHub, Microsoft, and OpenAI over GitHub Copilot and OpenAI Codex. These generative-AI tools were trained on public GitHub repositories and produce source code in response to users’ prompts. The programmers alleged that the tools sometimes reproduce code without the author names, copyright notices, license terms, and other attribution information that accompanied the original repositories.
The programmers relied on section 1202(b) of the Digital Millennium Copyright Act (DMCA). That provision prohibits intentionally removing or altering copyright management information (CMI), and distributing works while knowing that CMI was removed or altered without authority. After multiple amended complaints, the district court dismissed the DMCA claim with prejudice but allowed contract claims to continue. It certified the DMCA question for an immediate interlocutory appeal.
The Court’s Holding
The Ninth Circuit affirmed. It first held that the programmers had Article III standing at the pleading stage. Their allegations—including examples of Copilot reproducing portions of their code and research suggesting that large language models can emit memorized training data—plausibly showed a substantial risk that their code would be produced without attribution.
Standing did not save the DMCA claim. Judge Eric D. Miller’s opinion explained that “remove” and “alter” require an affirmative act directed at CMI connected to an existing work. Creating a new work that never contained the plaintiff’s CMI is different from copying an existing work and stripping its attribution. The court declined to impose a rigid rule that the original and challenged work must be literally identical. Minor changes cannot automatically evade the statute, and substantial reproduction without CMI may support an inference that attribution was removed. But the complaint itself described Copilot as learning statistical patterns and generating new code, not retrieving a stored copy and deleting attached CMI.
The court emphasized the boundary between DMCA claims and ordinary copyright infringement. Output that is substantially similar to protected code might support a traditional infringement claim, an issue the panel did not decide. Similarity alone, however, does not establish that CMI was removed from a copy. Treating every unattributed derivative or similar work as a section 1202 violation would displace ordinary copyright rules and expose defendants to the DMCA’s enhanced statutory damages.
The programmers also proposed an “input” theory: that defendants removed CMI before feeding code into the model’s training data. The Ninth Circuit did not decide whether that theory could state a claim because the programmers had not preserved it in the district court.
Key Takeaways
- A plaintiff can have standing to challenge a plausible risk that an AI system will reproduce its work without attribution, even when the DMCA claim ultimately fails on the merits.
- Section 1202(b) requires removal or alteration of CMI from an existing work or copy; merely generating a new, similar work without attribution is not enough.
- The Ninth Circuit rejected a strict literal-identicality test, leaving room for DMCA liability when a defendant substantially reproduces a work, makes minor changes, and strips its CMI.
- AI-training theories must be clearly pleaded and preserved. The court treated the alleged removal of CMI from training inputs as forfeited and expressed no view on its merits.
Why It Matters
This is an important appellate ruling on how a copyright statute written before modern generative AI applies to model outputs. For AI developers, it draws a meaningful distinction between generating code based on learned patterns and retrieving a copy whose attribution has been removed. For authors and open-source developers, it shows that unattributed AI output may require a conventional infringement or contract theory unless the facts connect the missing CMI to an existing copied work.
The decision is narrower than a general ruling that AI-generated code is lawful. It does not resolve whether training on licensed repositories infringes copyright, whether particular outputs are substantially similar, or whether removing attribution from training inputs violates the DMCA. Those issues remain open.
Surfaced via Law360 IP.
Your browser cannot display this PDF inline.
Download the full opinion (PDF)