Yes, since we complained about them violating copyright so much, it does not violate copyright if they destroy the original, so they are now destroying the original. We got our demands met.
Yay capitalism! Imagine how much value is being created for shareholders by this destruction! Another example of "great work" being done!
Jokes aside, I don't know what I would do in this situation either. Maybe find a way to mark "rare" or "out of print" books and sell/auction the "pages"? Or make the scanned book available via some mechanism?
On the surface it seems bad, although many of these books were probably rotting in place rather than being read. If the companies are willing to make the digital version available this may actually prreserve the books.
More troubling to me is the focus on pre-2022 data. This indicates that the companies are seriously worried about model collapse, where the future models become worse becuase they are being fed by generated data instead of real human data.
This is actually my worse case scenario for AI: the models become good enough to disrupt entire industries in the next 10 years, but then stay frozen at that level. Model drift then kicks in because the world will keep changing but the models do not, and in a few decades we actually regress instead of progress becuase they won't be enough skilled humans to drive advances.
Humanity hasn't caught up to the reality of the times; we should adjust copyright laws so anything over say 10 years old can be hosted for free on the Internet. We should have colossal databases of music, books, movies, etc. that can be freely traded and downloaded without breaking any laws. Maybe in the past the commerical protection of art and IP was necessary but with the Internet and AI I think we need to, as a species, get over it and focus on information archival and dissemination over protecting income streams.
I'm not really the most AI-friendly person around but this news cycle is just as equally annoying.
"Rare and out-of-print" is fabricating a lot of aura here. It's technically correct (the best kind of correct) but I've yet seen evidence that these are culturally significant copies being destroyed for scanning.
> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).
The oldest is from 1999! If that's the best they can actually enumerate to bolster this outrage farming cycle you could just wonder how irrelevant the rest are.
Really, please, kill this news cycle. There's a lot of issues deserving proper attention right now and this one is a straight-up nothingburger.
When Google scanned its books, it did not destroy any of the books. This was because of the idea of the "cultural object" where the book has value as a physical artifact. The experience of reading a physical book is different than reading an electronic text. There is an argument that reading a book in its original form is better because it is closer to the original experience. This can also be said of records and tapes. The experience is designed to work in the original format. A book is a designed object with cultural significance. The container for the words can be as important as the actual words. There should be clearer distinctions and policies between mass market production where items can be destroyed and there will still be many of them left and unique and original content with limited copies available. This is not hard to do. Not to do it shows carelessness.
17 comments
[ 0.21 ms ] story [ 14.1 ms ] threadJokes aside, I don't know what I would do in this situation either. Maybe find a way to mark "rare" or "out of print" books and sell/auction the "pages"? Or make the scanned book available via some mechanism?
Roughly nobody would have had access to that copy. Roughly everyone can now benefit from its content.
Confuses "(not all that) old books" with "rare books".
Ragebait, nothing more.
More troubling to me is the focus on pre-2022 data. This indicates that the companies are seriously worried about model collapse, where the future models become worse becuase they are being fed by generated data instead of real human data.
This is actually my worse case scenario for AI: the models become good enough to disrupt entire industries in the next 10 years, but then stay frozen at that level. Model drift then kicks in because the world will keep changing but the models do not, and in a few decades we actually regress instead of progress becuase they won't be enough skilled humans to drive advances.
https://news.ycombinator.com/item?id=49127284
"Rare and out-of-print" is fabricating a lot of aura here. It's technically correct (the best kind of correct) but I've yet seen evidence that these are culturally significant copies being destroyed for scanning.
From https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi...:
> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).
The oldest is from 1999! If that's the best they can actually enumerate to bolster this outrage farming cycle you could just wonder how irrelevant the rest are.
Really, please, kill this news cycle. There's a lot of issues deserving proper attention right now and this one is a straight-up nothingburger.