
On September 25, an article published by The Atlantic, titled “These 183,000 Books Are Fueling the Biggest Fight in Publishing and Tech,” caught my eye. Toward the end of the summer, Alex Reisner, the writer of that piece, had published a series of articles that revealed that a data set of books, called Books3, was among those being used by companies to train their generative artificial intelligence systems. In order to sound like humans, these systems need to be fed text written by humans.
The authors of the books included in this data set received no compensation and no royalties for the use of copyrighted works. As Reisner writes: “These authors spent years thinking, researching, imagining, and writing, and had no idea that their books were being used to train machines that could one day replace them. Meanwhile, the people building and training these machines stand to profit enormously.” Several writers in the United States have launched lawsuits claiming that this amounts to copyright infringement.