There is still no definitive legal answer as to whether training models such as ChatGPT, Gemini, and Claude on books and published works constitutes copyright infringement. The issue is not only whether the works were used without their owners’ permission, but also how they were obtained, the purpose for which they were used, and whether the output competes with the original market or changes the nature of the work sufficiently.
The issue is especially sensitive because many authors contributed, without their knowledge or consent, to the development of AI tools that may compete with their sources of income. Yet U.S. law, which has not been updated since 1976, leaves courts to apply old principles to technology that processes trillions of words and produces new texts.
The Anthropic Ruling Does Not Mean Training Is Illegal
In one of the most prominent rulings of its kind, Judge William Alsup ordered Anthropic to pay a massive settlement of $1.5 billion to a group of authors whose works were used to train the company’s models. At first glance, the ruling may appear to be a victory for authors, but the judge simultaneously ruled that the model-training process itself was lawful.
The penalty was connected to Anthropic obtaining the books from illegal online libraries, also known as shadow libraries, rather than simply using the works for training. Alsup compared the way language models absorb texts to the way a writer reads and studies literature and then produces a different work, rather than copying or literally replacing the original work.
Attorney Cathy Gellis, who specializes in intellectual property, copyright, and technology, believes this part of the ruling may be more useful to AI companies than to authors. The law focuses on copying and does not necessarily treat using, reading, or absorbing a work the same way it treats reproducing it.
Fair Use Does Not Resolve Every Case
Many cases revolve around the principle of “fair use,” an exception that allows copyrighted material to be used without explicit permission in situations such as criticism, parody, education, and others. Courts typically examine several factors, including the nature and purpose of the use, the amount of material used, and the effect of the use on the market.
Applying these factors to AI training, however, remains inconsistent. Jason Henderson, the founding attorney of the intellectual property and media practice at JWL International, explains that courts tend to be stricter when the purpose of training is to create a product that directly competes with the content owner or its market. By contrast, when the use does not appear competitive, courts may find ways to deem it permissible.
This distinction appears in the case of Thomson Reuters v. Ross Intelligence. Ross used Thomson Reuters content to build an AI-powered legal platform that competed with its product. Judge Stephanos Bibas concluded that Ross’s use was not transformative because it did not have an “additional purpose or different character” from Thomson Reuters’ content. Therefore, the use was not considered fair in that case.
As for the argument that chatbots compete with authors because they use their works to produce new artificial books, it has not yet achieved the same result before the courts, according to the article. This does not mean the argument has been definitively rejected; rather, the dispute remains open to different facts and rulings.
Training Is Not the Same as Rights in the Resulting Work
It is important to separate the question of using works to train a model from the question of whether texts produced by AI can receive copyright protection. In Thaler v. Perlmutter, the court ruled that a work produced entirely by AI does not enjoy copyright protection.
This raises independent questions: How can it be established whether a work was created by AI? And how can the proportion of the tool’s contribution be determined compared with that of the human author? Gellis uses a simpler example for comparison: a writer who writes a novel in Microsoft Word and uses its spell-check feature does not expect the company to become the owner of the novel. However, AI tools require a reconsideration of the boundaries between technical assistance and human creation.
certi.news Analysis: What Is Changing in Practice?
The real development is not the issuance of a single rule that permits or prohibits training, but the dispute’s movement toward more specific questions. The source of the data may change the legal outcome, as the Anthropic ruling showed, and the commercial purpose may be decisive when training leads to a product that competes with the content owner, as in the Ross Intelligence case. Therefore, it is not enough for companies to say that the model produces new text, nor is it enough for authors merely to establish that their works entered the training data.
For AI companies, the initial rulings mean that the procedures used to obtain training data have become a legal factor no less important than the design of the model itself. For authors and publishers, the current ruling does not establish an automatic right to prevent every training use, but it shows that the source of the copies and the final product’s effect on the market may provide a stronger basis for a claim.
The larger open question concerns conflicts among future rulings. Most AI companies still face pending lawsuits, and the initial decisions may affect the industry before they are later resolved or overturned in other courts. These cases therefore do not yet provide a stable solution; instead, they draw temporary boundaries that will remain subject to change as litigation continues.