October 7, 2026

A federal court has ruled that using copyrighted books to train artificial intelligence systems qualifies as fair use under United States copyright law. The decision, handed down in a long-running lawsuit against Meta, marks a significant early victory for AI developers who rely on vast quantities of published material to build their models. According to reporting from Futurism, the court determined that the act of training large language models on licensed and publicly available texts does not infringe on the rights of the original authors when the final output does not reproduce protected expression.

The case centered on authors including Sarah Silverman, Richard Kadrey, and Christopher Golden, who accused Meta of unlawfully copying their books to train its Llama language models. The plaintiffs argued that every time Meta’s system processed their works during training, it committed copyright infringement. They sought to hold the company responsible for both the ingestion of the material and the potential for the model to generate outputs that might echo their writing styles or specific passages.

United States District Judge Vince Chhabria rejected the core of that argument. In his opinion, the judge found that the training process itself constitutes a transformative use. Rather than reproducing the books for human readers, the AI system analyzes patterns, relationships between words, and statistical structures across millions of texts. This process creates an entirely new tool capable of generating novel text, summaries, code, and answers to questions. The court compared the practice to how humans learn by reading widely, noting that copyright law has never prevented people from internalizing ideas and language from books they purchase or borrow.

The ruling draws on established fair use precedents, particularly the Google Books case in which the Second Circuit held that scanning millions of books to create a searchable index qualified as fair use. In that matter, the court emphasized that the new purpose—full-text search—added substantial value without substituting for the original works in the marketplace. Similarly here, the judge concluded that training data serves a different function from the expressive purpose of the novels and nonfiction titles involved. The AI models do not sell copies of the books or allow users to read them in full. Instead, they produce original responses based on generalized patterns learned during training.

Meta welcomed the decision, describing it as consistent with how machine learning has operated for decades. The company pointed out that its Llama models were trained on a mixture of licensed datasets, public domain material, and web-scraped content that included books obtained through legitimate channels. Although the precise composition of training corpora remains largely secret for competitive reasons, court filings revealed that Meta had used datasets such as Books3 and other collections known to contain copyrighted works.

For the authors, the outcome represents a setback but not necessarily the end of the road. Their lawyers have already signaled plans to appeal, arguing that the ruling fails to account for the economic harm that may result when AI systems can generate works that compete with human authors in the marketplace. They also contend that even if training itself is fair use, the models’ ability to output material that closely mimics protected expression creates derivative works that infringe. The court left open some of these questions, noting that specific claims about infringing outputs would require separate evidence showing that the model reproduced protected elements rather than simply reflecting common literary tropes or styles.

Legal observers see the decision as part of a broader pattern emerging in AI copyright litigation. Similar lawsuits against OpenAI, Stability AI, and Anthropic have produced mixed preliminary results. In some cases, courts have allowed claims to proceed past the motion-to-dismiss stage when plaintiffs could show that outputs closely replicated their works. In others, judges have expressed skepticism that ingesting copyrighted material for statistical analysis alone violates the law. The Meta ruling adds weight to the argument that the training phase enjoys strong fair use protection, potentially influencing how other district courts approach parallel cases.

The decision also highlights ongoing tensions between different interpretations of fair use in the digital age. On one side stand technology companies that view large-scale data analysis as essential to progress in artificial intelligence. They argue that requiring licenses for every book, article, and website used in training would make development prohibitively expensive and slow innovation to a crawl. Many researchers point out that the internet itself was built on the ability to crawl and index content without seeking permission for every page. AI training, they maintain, represents a natural extension of that principle.

On the other side are creators who worry that their livelihoods are being undermined. Novelists, journalists, and visual artists have watched generative tools produce work that mimics their individual voices and styles with increasing accuracy. Some have seen their income decline as clients turn to AI services for first drafts, illustrations, or marketing copy. Professional organizations including the Authors Guild have called for new legislation that would require AI companies to disclose training data and compensate rights holders through collective licensing schemes.

The court’s opinion acknowledges these concerns but maintains that copyright law, as currently written, focuses on protecting specific expression rather than ideas, styles, or statistical patterns. Judge Chhabria suggested that if society wants to address the broader economic impacts of AI on creative professions, Congress should consider targeted reforms rather than stretching existing infringement doctrines. He noted that many of the plaintiffs’ strongest arguments sounded more like policy critiques than violations of the Copyright Act.

Beyond the immediate parties, the ruling could shape how AI companies approach data acquisition going forward. Firms that previously operated under the assumption that training on copyrighted material carried significant legal risk may now feel more confident expanding their datasets. At the same time, the decision is not a blanket immunity. The court stressed that fair use depends heavily on context. Uses that result in the regurgitation of substantial portions of copyrighted texts could still trigger liability. Companies will likely continue investing in techniques such as reinforcement learning from human feedback, output filtering, and retrieval-augmented generation to reduce the chance that models directly reproduce training material.

The case also raises questions about the future of licensing markets for AI training data. Several major publishers have already struck deals with AI developers, offering access to their catalogs in exchange for payment. If courts continue to view training as fair use, the incentive to negotiate such licenses may diminish unless companies seek them for public relations reasons or to gain access to higher-quality curated datasets. Smaller publishers and independent authors, who lack the leverage to secure such deals, could find themselves further disadvantaged.

Public discourse around the decision has been sharply divided. Supporters of the ruling argue that restricting training data would hand an advantage to companies with proprietary datasets or those willing to operate outside the law in jurisdictions with weaker copyright enforcement. They point to the rapid progress in language understanding, scientific discovery, and accessibility tools made possible by models trained on diverse public material. Critics counter that this progress should not come at the expense of the very creators whose work made it possible. They draw parallels to earlier technological disruptions in music and film, where new business models eventually emerged to compensate rights holders even as fair use protected certain transformative applications.

As the litigation moves forward, several related cases will test the boundaries of this decision. The New York Times lawsuit against OpenAI and Microsoft, for example, focuses heavily on the models’ tendency to output near-verbatim excerpts of news articles when prompted in specific ways. That case may turn less on the training question and more on whether the outputs themselves infringe. Visual artists suing Stability AI and Midjourney similarly emphasize how image generators can reproduce distinctive artistic styles, raising questions about whether style itself receives any legal protection.

The Meta ruling nevertheless provides a measure of clarity in an area that has been clouded by uncertainty. By affirming that the ingestion and analysis of copyrighted texts for training purposes can qualify as fair use, the court has given AI developers a stronger foundation on which to build. At the same time, it leaves room for accountability when systems produce infringing outputs or when companies engage in clearly unauthorized scraping of pirated material.

For authors and other creators, the path ahead likely involves continued advocacy for legislative solutions. Proposals under discussion in Congress include requirements for transparency in training data, opt-out mechanisms similar to those in the European Union’s AI Act, and the creation of compulsory licensing frameworks. Whether such measures gain traction will depend on the balance of political forces and the public’s perception of AI’s benefits versus its costs to creative workers.

The decision arrives at a moment when generative AI has moved from laboratory curiosity to everyday tool. Millions of people now use large language models to draft emails, summarize research papers, write code, and brainstorm creative projects. The underlying systems that power these experiences were trained on decades of human culture captured in books, articles, websites, and code repositories. How society allocates the value created by these models remains one of the central questions of the technological era.

While this particular lawsuit has reached a preliminary resolution favoring the AI developer, the larger debate shows no signs of ending. Appeals will likely reach higher courts, and new cases with different facts will continue to test the limits of fair use. For now, the court’s finding offers a practical guideline: training AI on copyrighted material, when done for the purpose of creating a new, transformative tool rather than reproducing existing works, falls within the bounds of what copyright law permits. The challenge for lawmakers, technologists, and creators will be to build on that foundation in ways that encourage innovation while protecting the human creativity that makes such innovation possible.

Federal Court Rules AI Training on Copyrighted Books Is Fair Use first appeared on Web and IT News.

Leave a Reply

Your email address will not be published. Required fields are marked *