Artificial intelligence companies are buying millions of physical books, cutting them apart, scanning every page and recycling the remains to train the chatbots powering today's AI boom, according to a report by The Washington Post, based on newly unsealed court filings in a copyright lawsuit against AI startup Anthropic. The filings also reveal how the race to build more capable AI models pushed companies to secure vast collections of books, setting off a wave of copyright lawsuits.
The court documents provide one of the clearest glimpses yet into the AI industry's search for high-quality training data. While Anthropic's internal book-scanning project is at the centre of the filings, separate lawsuits involving Meta, OpenAI and Google also highlight how leading AI companies sought access to millions of books to improve their models, although the methods differed across companies.