Get all your news in one place.
100's of premium titles.
One app.
Start reading
The Conversation
The Conversation
Agata Mrva-Montoya, Senior Lecturer, Department of Media and Communications, University of Sydney

Meta allegedly used pirated books to train AI. Australian authors have objected, but US courts may decide if this is ‘fair use’

Companies developing AI models, such as OpenAI and Meta, train their systems on enormous datasets. These consist of text from newspapers, books (often sourced from unauthorised repositories), academic publications and various internet sources. The material includes works that are copyrighted.

The Atlantic magazine recently alleged Meta, parent company of Facebook and Instagram, had used LibGen, an illegal book repository, to train its generative AI tool. Created around 2008 by Russian scientists, LibGen hosts more than 7.5 million books and 81 million research papers, making it one of the largest online libraries of pirated work in the world.

The practice of training AI on copyrighted material has sparked intense legal debates and raised serious concerns among writers and publishers, who face the risk of their work being devalued or replaced.

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.