An AI company has “destroyed” millions of books as part of a “controversial project” to train its models, according to City A.M.
Tech giants want to “feed an insatiable demand for training data”, but their methods could have serious ramifications for the publishing industry and the quality of future books.
The alarm was first raised when antiquarian booksellers in the UK and across Europe started to receive curious requests for thousands of obscure titles from anonymous buyers.
According to recent court documents, Anthropic, the developer behind Claude, launched an internal programme known as “Project Panama”, which concluded that books were “essential” for training advanced AI models.
One senior executive at the AI giant said books taught systems “how to write well”, as opposed to using “low-quality internet speak”. The company has bought millions of second-hand books, “sliced off their spines” using “industrial cutting machines”, scanned every page and recycled the rest.
This trend has “raised fresh concerns about copyright and the future of physical books”, said the Turkish news outlet Anadolu Agency. But dealing with the issue could be quite another matter; copyright violations are “difficult to prove in court” because of “legal loopholes”.
Thomas Koch, of the German Publishers and Booksellers Association, said it “appears to be yet another example of AI companies using vast quantities of copyright-protected works to train their language models, without consent and without payment”.
|