Authors Are Furious After Finding Their Works on List of Books Used To Train AI

Authors using a new tool to search a list of 183,000 books used to train AI are furious to find their works on the list.

https://www.themarysue.com/authors-are-furious-after-finding-their-works-on-huge-list-of-books-used-to-train-ai/

Authors Are Furious After Finding Their Works on List of Books Used To Train AI

Authors using a new tool to search a list of 183,000 books used to train AI are furious to find their works on the list.

The Mary Sue
Here’s an idea, legally force companies like OpenAI to rely on opt-in data, rather then build their entire company on stealing massive amounts of data. Sam Altman was crying for regulations for scary AI, right?
Would search engines only be allowed to show search results for sources that had opted in? They "train" their search engine on public data too, after all.

They aren’t reselling their information, they’re linking you to the source which then the website decides what to do with your traffic. Which they usually want your traffic, that’s the point of a public site.

That’s like trying to say it’s bad to point to where a book store is so someone can buy from it. Whereas the LLM is stealing from that bookstore and selling it to you in a back alley.

AI isn’t either. It’s selling statistical data about the books.
“I’m not reselling your book, I am selling a machine that holds a mathematical formula that partly represents your entire book word for word and can reprint it on command!”
LLMs can't reprint their entire training data on demand. They rarely even remember quotes.