I don't remember what thread it was in about AI using copyrighted works for training. But I just read this:
A U.S. court approved the largest copyright recovery in U.S. history, penalizing Anthropic $1.5 billion for building its AI with pirated books.
In the largest known copyright recovery in U.S. history, a federal judge finalized a historic $1.5 billion class-action settlement between artificial intelligence company Anthropic and a coalition of authors and publishers.
The legal battle began when thriller novelist Andrea Bartz and other creators accused the Claude chatbot developer of downloading hundreds of thousands of copyrighted books from online piracy sites like Library Genesis. Rather than face a lengthy trial, Anthropic agreed to the massive payout, which covers roughly 482,000 books. Eligible rights holders are set to receive about $3,000 per affected work, with over 91% of the eligible works already claimed by authors and publishers.
While the record-breaking settlement marks a monumental victory for writers, it leaves the central legal debate surrounding generative AI training unresolved. The court drew a crucial distinction, noting that while the technical training of AI models might qualify as fair use, the unauthorized acquisition and storage of pirated materials is a clear copyright violation. This ruling indicates that while AI firms can legally learn from copyrighted texts under fair-use guidelines, they must procure their training data through authorized, paid means rather than shadow libraries. As dozens of other copyright lawsuits loom, this historic decision signals that the fight over creative rights in the AI era is far from over.
source: Associated Press. (2026, July 21). Judge approves a $1.5B Anthropic settlement over books used to train Claude. AP News.