Authors Are Furious After Finding Their Works on List of Books Used To Train AI
Authors using a new tool to search a list of 183,000 books used to train AI are furious to find their works on the list.
Authors Are Furious After Finding Their Works on List of Books Used To Train AI
Authors using a new tool to search a list of 183,000 books used to train AI are furious to find their works on the list.
They aren’t reselling their information, they’re linking you to the source which then the website decides what to do with your traffic. Which they usually want your traffic, that’s the point of a public site.
That’s like trying to say it’s bad to point to where a book store is so someone can buy from it. Whereas the LLM is stealing from that bookstore and selling it to you in a back alley.
It shares popular quotes from books, it can’t reproduce arbitrary content from a book. The content needs to be heavily duplicated in the training data to stick around, and even than half of it might still end up being made up on the spot.
Also request for copyrighted content will be blocked by ChatGPT and just receive the stock “I can’d do that” response anyway.
If you have some damning examples that show the opposite, show them.