This is a good thing to a point, right? They sell heaps of stock, some old / outdated stock (like "Pass Your Driving Test, 2018 Edition" as per the article), in bulk, at full price, no questions asked. Retailer's dreams come true.
(I don't believe them being digitized for consumption by AI will impact book sales much, as ebooks, google books, project gutenberg, etc didn't either)
Is this a second form of AI psychosis, as in AI training corpus psychosis?
The need for all the content, Moar!! Feed me. Is this the Paperclip Maximiser in the form of the Training Material Maximiser?
Will it be that, in the end, the lack of "Pass Your Driving Test, 2018 Edition" was the cause of driving rule hallucinations in all previous models? Will this finally get OpenAI back in front of Anthropic? Hurrah! We found it!
> with a focus only on books with an ISBN number, the identification system introduced in 1966
So it's not likely to be rare and precious books. It's things which are already in the Library of Congress or its many equivalents.
If they're forced by laws to destroy the results of the scans after training on them, as some have implied, that's bad, on the chance there is some actual lost media in there. But if they keep the scans, or even the transcripts, that's probably an improvement on the status quo, to be honest.
I've had family members in the used books business, and trust me the fate of the vast majority of these books was always to be pulped.
Or have their covers stripped off and shipped out as "remainders."
Almost all the 60s science fiction I've owned came like this, bought on the pavement in port cities like Bombay and Madras. Container loads sold by weight.
I didn't think about this happening. It reminds me of the cross-border challenges with crypto. Each country has laws to control its author rights or money supply, but those are hard to enforce in the international setup.
When I hear these stories of AI companies buying all the books, I think back to Kevin Kelly in 2008 or 2009. He talked about how books are less expensive than at any other time in history and easier to buy and that could change so it makes sense to buy a lot of books.
> [...] I was near to the point of actually digitizing and getting rid of all my paper books.
> I was that close about five years ago, but then I had an
epiphany. I went to private library, and I realized that books
were never as cheap as they are today. They never will be as
cheap, and that there's some power about having these things in
paper always available, no batteries, never obsolete, and that if
you made a library now, you would never be able to make some
of these libraries in 50 years, so I decided to keep and to
cultivate this paper library as something that was going to be
powerful in the future.
I remember seeing photos of his library but can't find them anymore. This is the only thing I could find:
I wish we had competent regulators. Force the scans to be purchased through an official repository. Save the scans and sell them to other companies who also want to train models. Force the companies to make all book requests through the official repository. Give some portion of money paid back to the publishers. This is a solvable problem.
Two American self published nonfiction authors I know reported large purchases of new books earlier this year, more than 50 units at a time, via Amazon. This is unusual because the books are not well known. Because the books do not sell well otherwise, they were very few used copies available for sale on Amazon.
Here's my AI theory: someone is purchasing lots of books en masse to be packaged and sold for LLM training to multiple clients (not just a single company like Anthropic), but only scanning once and then reselling the scans while preserving a "chain of custody" proof that individual copies were purchased. Claude shot that idea down for reasons related to the intricacies of US first sale doctrine and copyright law, but had an ambiguous response when I proposed that it might be Chinese companies doing something similar for training AI models that are intended for possible resale to overseas markets.
21 comments
[ 0.28 ms ] story [ 12.6 ms ] thread(I don't believe them being digitized for consumption by AI will impact book sales much, as ebooks, google books, project gutenberg, etc didn't either)
The need for all the content, Moar!! Feed me. Is this the Paperclip Maximiser in the form of the Training Material Maximiser?
Will it be that, in the end, the lack of "Pass Your Driving Test, 2018 Edition" was the cause of driving rule hallucinations in all previous models? Will this finally get OpenAI back in front of Anthropic? Hurrah! We found it!
So it's not likely to be rare and precious books. It's things which are already in the Library of Congress or its many equivalents.
If they're forced by laws to destroy the results of the scans after training on them, as some have implied, that's bad, on the chance there is some actual lost media in there. But if they keep the scans, or even the transcripts, that's probably an improvement on the status quo, to be honest.
I've had family members in the used books business, and trust me the fate of the vast majority of these books was always to be pulped.
Or have their covers stripped off and shipped out as "remainders."
Almost all the 60s science fiction I've owned came like this, bought on the pavement in port cities like Bombay and Madras. Container loads sold by weight.
why not just every textbook for every subject in high school and college?
Excuse me, what?
Public domain books - no need to bother, they’ve scraped PG and any other source and can do with it what they like.
Used booksellers should raise the price of used books that are of no interest to any real reader.
> [...] I was near to the point of actually digitizing and getting rid of all my paper books.
> I was that close about five years ago, but then I had an epiphany. I went to private library, and I realized that books were never as cheap as they are today. They never will be as cheap, and that there's some power about having these things in paper always available, no batteries, never obsolete, and that if you made a library now, you would never be able to make some of these libraries in 50 years, so I decided to keep and to cultivate this paper library as something that was going to be powerful in the future.
I remember seeing photos of his library but can't find them anymore. This is the only thing I could find:
https://colossus.com/article/flounder-mode/
What's the use of all this AI if they can't figure out this relatively simple task.
Here's my AI theory: someone is purchasing lots of books en masse to be packaged and sold for LLM training to multiple clients (not just a single company like Anthropic), but only scanning once and then reselling the scans while preserving a "chain of custody" proof that individual copies were purchased. Claude shot that idea down for reasons related to the intricacies of US first sale doctrine and copyright law, but had an ambiguous response when I proposed that it might be Chinese companies doing something similar for training AI models that are intended for possible resale to overseas markets.
AI companies are shredding rare books
https://news.ycombinator.com/item?id=49068738