I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
The articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
At this scale, there are no guarantees of anything. There very likely will be unique copies in there. If these were expert archivists a lot of damage could be prevented, but given the malice and indifference of AI companies, there very likely will not be an expert archivist involved, and unique copies will be destroyed unceremoniously.
The 404 story suggested that these are largely vanity press books and instruction manuals for things no longer sold. Implying that these books were almost certainly headed for the recycling center had the AI companies not snatched them up.
Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? Hmm.
Books by Zolar are interesting hard to find all editions. The Fearful Void by Geoffrey Moorhouse probably still has 100s of copies available but hate to see it lost.
The hypothetical "heroic figure goes and buys last copy of a 1962 guide to Ford cars to carefully maintain it in an appropriately climate controlled library" is vanishingly unlikely. A ten or a hundred or a thousand times to one, it just goes to the trash. At least here it gets scanned by the AI company.
Dumpsters also don't typically come equipped with a robot scanner and network uplink built in.
Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.
That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.
> I really don't know what people objecting to this imagine typically happens to old, unwanted books.
In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops.
Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there.
95 years after publication. Many other countries also have an X years after the author's death clause where X varies between countries but is at least 70.
There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further.
The thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, when it's free to share.
At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.
Since they seem to leapfrog each others’ models every few months, the training data is one of the few ways they can build competitive advantage, and that explains why they don’t share, even if we don’t have to like this.
This is solved by law, which is solved by 'we the people' and I bet many AI companies would be fine with something like the equivalent to patent law with bankruptcy escrow to the library of congress, where they must release the scans in 10 years for books that the vast majority will not give a flying shit about. By then the advantage is long gone in data moat.
In my opinion this is one of the reasons why libraries should accept any book, even if all they do is examine it and throw it in the trash. This way they would have a chance at finding any treasures that could be regularly dumped in that way.
> They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies?
I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.
Many are just copyright "orphans", nobody knows who owns the copyright any longer. Maybe the author died and the copyright passed to their estate, but they're not even aware of it.
One book I'm hunting for a copy of right now was published in England in 1947 and in those days paper was rationed, so not many copies were made, and only a handful have survived. As soon as I find it I'll scan it and upload it to IA.
So, have you tried finding out what the programming was in October 1994? Or what cultural ephemera appeared in the TV guides of that era alongside the schedules? Either there's a copy for the week you want in an archive, or somebody's got one for sale, or most often neither. This can piss you off, if as it happened you had a reason to care.
To play devils advocate, completely on the terms of your argument, would it be better for that particular human artifact to be shredded and its contents melted into an anonymized data pool, or for it to exist in a museum archive, in its original form, such that future generations can better understand what it was like to be alive in 1994?
I’d personally choose the latter, especially given that the 1994 tv guide is not going to meaningfully improve the utility of the language models.
Direct access to pre-digital history is drying up rapidly, why accelerate that for incremental benchmark gains in a domain that isn’t even relevant to the most useful forms of a nascent technology?
A museum - or any other building - can only hold so much physical stuff. How much of it do you really want preserved? How do you choose what is preserved (it's an eventually inevitable choice)? Do you save the 1980s stuff but not the 90s? Or save every even/odd year? Some other method? How much direct access do you think people need to pre-digital history?
A reasonable opinion, but I'd personally strongly choose the former.
A physical book in a museum archive is useless for 99% of the worlds population even if they really wanted that specific book and were able to find it, as they'd have to arrange for access, then travel (at incredible expense) to access it.
Maybe they could ask the museum to digitize it, but that's still going to be days of delay and tens of dollars of cost to access parts of that book, if the museum even offers that service. If we go slightly beyond your "melted into" statement, the chances of the book becoming useful to the public are much higher in the AI company's digital archive, which might turn into something like what Google Books could have been, given the right incentives and copyright law changes.
And of course that presumes that the TV guide is going to stay in the museum rather than been thrown out as part of curation (or realistically, long before it makes it into a museum). Neither museums nor archives hoard everything, throwing stuff out is - as far as I know - one of the key jobs of an archivist. And a 1994 TV guide, while useful to understand what it was like to be alive in 1994, likely doesn't contain much unique information. You don't need that specific guide.
If there are 52 weekly editions, of 10 different guides, you would likely get most of what you want from any one of them. And for the parts that you wouldn't - there's a good chance that you'll have a much easier time getting the essence of this knowledge from the anonymized data pool that all the content was melted into, rather than chasing 10 different museums to find the original magazines.
Isn’t this also counting on the AI company faithfully reproducing the contents of the book, and no hallucinations or shenanigans occurring? How will we ever know what the book actually said if people wanted to argue its contents later?
Well for example yours - if you don’t pay for these rare books and don’t store them in good condition, then you have decided that they are not worth saving. Many of these rare books would run you like I don’t know, 1 buck?
I don't think I have enough money or room to buy all the books I think are worth saving, and my list no doubt has some overlap with someone else's. If I did buy them, what if someone else wants to read them? If I lose it or it gets destroyed in some way, the chance it will be gone forever increases (assuming there are multiple copies). Entrusting the availability of knowledge to individuals like that sounds like a bad idea. There should be some kind of publicly funded organisation that can take care of a big collection of books, afford to keep them safe, and make them accessible to anyone.
Those are the rare books you come across in your life because of the way you choose to live, if you were literate and in the habit of reading literature, mathematics, science, you would realize that there are many texts that are at present exceedingly difficult or impossible to get your hands on
> Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
> ...AI companies will still do scan'n'destroy because it's just cheap.
There is also a legal element. If they kept the physical copy around after scanning the argument is that they're making copies of the book which puts them on tricky legal ground. By destroying the physical copy they can argue that there is only one version of the book that now exists solely in digital form, so this usage is better protected under fair use.
Yes. I dont understand at all what AA is worried about. One copy of a book is no big deal? good will and used book stores throw out a lot more than that.
I’m not aiming this at you directly by: ISBNs or STFU
Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this leads me to believe it’s people worked up about the idea of destroying (except it’s not destroying, it’s transforming, a fact often ignored) books, books that it’s not clear at all there is any strong demand for.
People want to invoke things like F451 but it doesn’t compare in the slightest. It’s like when people get mad about libraries throwing away or otherwise liquidating books that no one is reading in order to bring in books people want to read. People get all up in arms about that as if a book itself, in isolation, is inherently valuable or worth protecting. It’s not. If no one wants to read it then what value does it have? The impetus is on the people that think the book has value, it’s on them to carry the torch, to preserve what they think is worthy.
It would be like a company going to a yard sale and buying unsold/unwanted items to 3D scan them and destroy them in the process. This isn’t breaking into the Louvre and destroying one-of-a-kind artwork.
"A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
"But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
"It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
It would be if it made the point you think it is. All of this is more and more hand waving. 75 copies in museums? Then I think we’re good. As for the 18th century books, that’s pure speculation. It’s like a museum saying, “yes we have they prints for sale that they keep buying and destroying but wouldn’t it be a shame if someone destroyed the actual Mona Lisa?”.
And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them. This isn’t complicated. Amazon/etc aren’t breaking into museums and libraries, they are buying books on the open market.
If these books are so rare and important, then why has no one cared until now to actually preserve them?
> It would be if it made the point you think it is. All of this is more and more hand waving. 75 copies in museums? Then I think we’re good.
You're repeating the seller's contextual point, as if it's a counterargument. Do I need to explain to you that what makes the lone surviving 18th century edition important is that there aren't 75 copies of it in museums?
> And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them.
It's good to see you agree any digitization of this category of book should be non-destructive.
> If these books are so rare and important, then why has no one cared until now to actually preserve them?
(Lastly for realz this time, eh?) Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.
You are conflating 2 parts of the article to make it say something it's not. No one has put forward any examples of a "lone surviving 18th century edition" being bought up my AI companies and destroyed. That was just an example of "wouldn't it be terrible if", not a "this has actually happened". It's a complete hypothetical, it's a made up scenario, it's a boogeyman. You've hung your entire 2 comments on something that has no evidence of happening.
> It's good to see you agree any digitization of this category of book should be non-destructive.
I don't. Digital or physical, it's the same, there is no difference in my mind. I have many paper books but they are art, not functional, I also have the ebooks which is what I actually read. The _ideas_ are what's important, not a dusty, decaying shell in which the ideas are contained. I wouldn't shed a tear over every library digitizing their books and destroying the physical versions, nothing is lost. More importantly, once you've bought something it's yours, yours to read, yours to display, yours to destroy. On hacker news, of all places, the people fighting _against_ first sale doctrine is appalling.
> Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.
Our definitions probably differ here but preserving is not storing a book, preserving is ensuring that even if this copy is destroyed the ideas inside live on. I think that people that hoard (actually) rare books without a thought or care to making sure the text inside is preserved for future generations out of some desire to simply own something rare are the actually monsters here. And let's dispose with the notion that booksellers are "preserving" books, they are holding inventory, inventory they were happy to sell to Amazon/etc. If AI companies were raiding museums at gunpoint we'd be having a different discussion. They are buying books for sale, if they keep them in a library at corporate or scan and destroy them it makes no difference.
If you're buying second hadn books by the lot, you'll get a lot of duplicates and its eaiser to scan wholesale and dedupe in the computers than it is to try to run a sorting operataion on "things".
Sorry, is your stance seriously that authors and publishers should digitize and freely distribute their work, at their own expense?
Also, who’s forcing AI companies to “ingest” books in such a destructive way?
Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.
Relinquishing copyright does not imply any of the labor you're suggesting. It's the opposite: you're just committing not to perform the labor of pursuing legal action against someone who does digitize and freely distribute the work.
Anna's Archive, for one, would be more than happy to host at no cost to the author.
AI companies are buying the physical books, they can turn them into confetti if that's what they want to do. If the physical books are running out, the authors can print and sell more. Or they can sell digital copies so the information is not lost.
The law is currently forcing these companies to destroy the books after scanning them.
> The law is currently forcing these companies to destroy the books after scanning them.
No, the law is stopping them from digitally sharing their scans. They are perfectly capable of reselling or donating or storing the books they buy. (Wasn’t Amazon originally a book seller?)
> Sorry, is your stance seriously that authors and publishers should digitize and freely distribute their work, at their own expense?
digitize: no, there's no burden to do this
freely distribute: no, there's no burden to do this
use it or lose it on the copyright: yes; if a work is copyright but a publisher doesn't want to make new copies because they won't make money on it then it should be free to copy
They do it because, for each work, they bought one copy, which they scan and no longer need the physical version of, and would be in copyright violation if they keep more copies than they bought.
Copyright holders are certainly responsible for keeping unavailable works inaccessible, but shredding is mostly an industrial scanning decision, not a copyright requirement
I was wondering what the pro book shredding take was going to be. Why destroy the book after scanning? You can create a beautiful library of rare books with all the AI debt bubble.
It's also the regulation, where most systems still look at "one pirate copy" = "one sale of lost profits", especially when pirates end up in court. If the book (or game or whatever) is not sold anymore in any way where you could give the copyright holder money in an easy accessible way (eg. buy it on amazon, or a local bookstore), they shouldn't be able to claim losses from piracy, since they clearly don't want your money.
On the other hand, there are grey zones here, the lord of the rings books (still copyrighted and easily obtained pretty much everywhere) have been translated into my language many decades ago, and many of us read and liked those translations, but when the movies came out, a new translator did a new translation, where they changed a lot of things, including the last names of bilbo and frodo (Bogataj->Bisagin) and the Shire (Grofija->Šajerska), and the old version is sadly available only in paper form on second hand markets. On one hand, copying that if you only want this specific version would not cause a lost sale, on the other, you can get new translations (or english originals) pretty much everywhere.
Not the copyright holders, "we the people": Copyright is an artificial legal construct that was repeatedly ratcheted up over and over again.
Unfortunately 50 years after the death of the author (or 50 years after publication for corporate owned works) has been locked in as a minimum term through international treaties, so it'll be somewhat hard to lower it beyond that, but many countries (including the US) enforce much longer terms, so that would be a first lever that could be applied quickly.
Maybe countries could could also establish an exception for out of print books offered to the public for free, or a general "library exemption" for public archives after a certain number of years?
I'm sure one of the AI companies would be willing to host a LibGen style library as a PR measure if legally allowed (with sign up required for rate limiting and as an extra benefit for the company to get daily active users).
With many old books, a big part of the problem is that it's non-trivial to determine who owns the copyright. Sometimes the contract would say the copyright reverts to the author after a certain amount of time out of print,but you have to go dig through old contracts to figure out whether that's the case for any given book.
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
In the context of books, "scanning" is now more commonly something that should be called "camming" --- you simply point a camera at the book, and take a picture of every page.
Not everyone should scan stuff. If you spend any serious time looking through stuff that randos on the Internet have scanned the quality fits the Bell Curve perfectly.
Biggest problems:
- scanning items that are bigger than the scanner platten so the start/end of every line is cut off.
- becoming an "editor": scanning only the pages you think are interesting and skipping intros, forewords, title pages, copyright pages etc
I work in this space. I now require that before scanning a video is made carefully flicking through every page of the item so it can be checked after scanning to ensure all the pages are present and in the original order.
Even the big libraries fuck up. I wanted an intact copy of Harper's Weekly from 1900 that has a big fold-out map in it. It's not clear to the libraries scanning this issue that the map is missing from their copies. None of the copies for sale from dealers have the map. Even when it is still glued into the middle it gets missed by industrial scanners. Google's scan only includes the (blank) back of the folded map.
Luckily GPT was able to track down a copy in a university special collections and fired off an email asking them to scan it. I just got the scan today:
I spend a lot of tokens getting LLMs vision tools to find the missing pages in vintage items and then try to reassemble them from other scans where available.
I'm also splitting up volumes to reupload. A lot of periodicals are only available online as giant multi-gig volume PDFs with all the issues in one file. I have a separate app I wrote to scan all the pages looking for covers so they can be split into PDFs and then identifying the volume/issue/month/year data from the cover or title page.
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
There was no way in your country to pay and watch it? FIFA will have sold the tv rights there to someone, surely. In which case your complaint is what, that it was expensive?
Free on both, but not with the narration in my native language.
On Brazil the world cup was being transmitted on youtube. in NL only on traditional TV channels or Online for the same channels (all for free but in Dutch).
And literally as I write this I receive an email saying that my youtube premium was raised from 33 to 38 EURO. So there we have piracy getting juicier and juicier.
There is also a ton of tv shows, movies, music and books that you cannot buy, for now real good reason. I wouldn't be surprised that if in a few years there will be shows and movies that are only exists as pirated versions. With things increasingly only being available on streaming platforms or behind DRM in other ways, we risk looking back on the current era as a black hole 50 years from now.
My concern is that less popular content is just erases, lost in mergers or lost in massive datacenters, never to be seen again.
> I wouldn't be surprised that if in a few years there will be shows and movies that are only exists as pirated versions.
I can already think of a couple examples I've run into in the audiobook world. I have copies of Douglas Adams himself narrating his Hitchhiker's Guide to the Galaxy and subsequent books. As far as I know, these recordings are not available for purchase anywhere. I got them because someone was kind enough to upload them. I would have preferred to buy them but didn't have that choice.
Some of Iain Bank's audiobooks were only available in Europe for a time (and maybe still are). For those I was able to convince Amazon I was buying in Germany at least.
Indeed, art. And it is a pretty common position that all people should have access to art and culture. Add up the cost of buying the DVD/Blu-Ray releases for the 1500 or so films that make up the canon of cinema. That's a sum of money daunting even for people in developed countries, let alone most of the world. Piracy is going to be the realistic solution. (And before you say "Use the library", you know well-stocked libraries don't exist in most of the world, right?)
Even if not all art, once something becomes canonical, it then becomes something that people should be able to easily familiarize themselves with for the sake of an educated and edified citizenry. And indeed, many countries subsidize public libraries and live performances for this very reason. But no state can manage to provide free or nearly-free access to the entire canon, so piracy helps fill the gap.
If you're saying "piracy is justified in those parts of the world where access to any art or culture is prohibitively expensive", then that's a pretty defensible position, not unlike "It's ok to steal bread if you are starving".
But that's not what many people defending piracy are doing. Many of them do indeed have great libraries nearby, and art and culture on demand at prices they happily spend having food delivered to their house.
I don't know if you have noticed, but piracy has declined greatly since the introduction of streaming. The scene is a shadow of its former self. The communities today that are keeping high-quality releases of canonical music and films available are overwhelmingly based in regions of the world with a dearth of good libraries.
Anna's archive is selling data to AI companies. They're essentially saying "hey, don't sell your books to be scanned by AI companies, scan them yourself, and give us the data, so we can sell it to AI companies."
Why are the companies legally required to shred the books? That's the most surprising part about this to me. Surely if they bought them second hand they could donate or resell after scanning. I'm wondering if the scanning machines are damaging the books.
The legal idea behind it is that if you "copy" the book it's bad because now there are two copies and you "stole" from the author, but if you "move" the book to a digital form (and don't copy that outside of your organization), it's fine.
Given that they have to destroy them anyway, they're obviously also going to use the much cheaper destructive scanning (cutting off the spine and using a feed scanner rather than carefully turning page by page).
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
Maybe the law doesn't explicitly say that. But it's easier for lawyers for AI companies to argue the case if they destroyed it. "Look, there's only one copy! We didn't make additional copies!"
It's also probably more convenient for them to destroy the books as opposed to trying to find space to store them. Knowing those companies, most likely they'd be just stuffed into some warehouse to rot after a couple years.
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
More and more I feel like anti-AI is a bigger bubble than AI. It seems like every week it expands into a new dimension - anti-Flock protesters tearing down years-old traffic cameras that were used for research into auto accidents, etc.
Like, the current thing in the news cycle is a poll that young people are now more worried than hopeful about AI. Which sounds scary, but my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..
> my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..
There are were polls about people being "more worried than hopeful" about satanic cults or alien abductions? With the youth being the most worried about satanic cults? Just like that knee-jerk "it's the bigger bubble", that makes zero sense.
As per the GDC 2026 State of the Industry poll, "52% said gen AI is bad for the industry, nearly double the 30% who held that view last year". But sure, everybody but HN, LinkedIn, and X bros are just luddites clutching pearls in their tiny bubble. They're the weak and stupid ones, and that is why the stupid shit said about them, day in and day out, isn't actually stupid. It all checks out.
I don't know how it works today, but 20 years ago bookstores wouldn't return unsold books (too expensive to ship) but would simply tear the covers off and throw them in the garbage.
Its nuanced and complicated but its not fully without reason. I think there would be a compromise of using those books while playing the "we're helping preserve them part" but I don't think that even crosses the mind of most Ai CEOs in an honest way.
The crux of the issue is that it is mostly a PR problem. AI is amazing but being promoted, in the eyes of many, by the worst people imaginable. Very akin, and overlapping in many ways to the crypto crowd.
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
There's no "2022 split", and the "untouched by machines" bit came from the marketing blurb of a company offering book scanning services - not the AI labs themselves.
At the AI lab level: the book scanning seems to be driven by copyright concerns, not data contamination concerns. There was a concern about AI contamination, but there's no measurable performance loss from ingesting post-2022 data with minimal filtration, and some tests attribute small but persistent performance gains to post-2022 AI contamination. It's unclear where exactly do those gains come from.
Why is all high quality text useful? The "inverse problem" framing is that all text reflects the thinking behind it, somewhat, and by learning to reproduce it, LLMs implicitly learn to reproduce some of the thought process too. They don't just memorize the dry factual knowledge, but also learn how that knowledge fits together, and how to reason about that knowledge - both in the specific case and in general. And that "in general" then surfaces in an LLM's ability to generalize. Which is very desirable.
Also there was a highly discussed paper talking about how “touched by machines” content will kill llms. About a month after the papers first llms trained with “touched by machines” content appeared an the capabilities of the models got huge upgrade by using that dirty content
I entirely believe the litigation brought against Internet Archive was secretly sponsored by these exact organizations, because they want to monopolize information to train models.
> On March 24, 2020, following shutdowns caused by the COVID-19 pandemic, the Internet Archive opened the National Emergency Library, removing the waitlists used in Open Library and expanding access to these books for all readers. More than one user could borrow a book at the same time. Two months later, on June 1, the National Emergency Library (NEL) was met with a lawsuit from four book publishers. Two weeks after that, on June 16, the Internet Archive closed the NEL, and the prior Open Library CDL system resumed after the 12 weeks of NEL usage.
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
Public libraries destroy unsold book donations all the time. I often tried to give away some old books I have online and nobody wants them. Some of these books have some nostalgic value to me so I hate to see them just get destroyed so they just lie in my shed.
I have a copy of Michael Abrash's Graphics Programming Black Book (it's like 1k+ pages) with DESTROY written in red on the sides. I appreciate that someone saved it and sold it to me for cheap :)
Which is something I argue against constantly (and at least my local library tries hard not to) --- discarded books are placed on tables near the children's area usually and folks are free to pick them up (it might be that a few of them are sold, I certainly get a lot of ex-library books when buying on Thriftbooks and Better World Books).
A partial solution there is of course a larger budget and a "last copy" policy where the last copy of a text at least is stored away in deep storage against a future loan.
The local libraries also accept book donations for an annual fund-raising sale.
The question I have is, do these companies keep copies of the scans after they have finished training on them? If so, then it isn't the worst outcome. Not great but at least the information is not completely destroyed forever just the original physical being of it.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
A lot of Hathi is locked behind university and library access restrictions. I sometimes have to track down students or someone who has a local library card to get items I need.
What evidence do we have that they are "destroying" books?
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
I've read a few of those articles in recent months, both in English and German, and I did read any book title that was rare. The Rare Book & Special Collection Div at the Library Congress considers books published before 1801 as rare. Searching for old books on abebooks is surprisingly hard but I didn't find any for less than 10 Euro and nothing in those articles suggested they were buying up anything but cheap books.
I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!
The hysteria around AI and data centers has hit a precipice. It's actually a bit embarrassing now. I am pretty sure there are foreign adversaries that are trying to stop the US, but I also really blame the AI companies for doing the most horrendous job imaginable in pitching AI to the public. Not a shock that people are against something that tech bros have claimed will destroy everyone's lives in the next 5 years. These books were probably going into a landfill without AI companies getting them, regardless. Tons and tons of books go into the garbage every day.
I can imagine 100y from now, most if not all books and knowledge are in electronic format or even just as part of an AI, then a wild solar flare wipes out all electronics in a minute..
Books have been declared dead several times of the last quarter century, yet in the EU alone it's still half a million new titles every year and a 40 billion Euro market.
I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!
230 comments
[ 71.6 ms ] story [ 2261 ms ] threadInstead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.
That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.
In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops.
Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there.
There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further.
In short, it's a mess.
At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.
I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.
One book I'm hunting for a copy of right now was published in England in 1947 and in those days paper was rationed, so not many copies were made, and only a handful have survived. As soon as I find it I'll scan it and upload it to IA.
I’d personally choose the latter, especially given that the 1994 tv guide is not going to meaningfully improve the utility of the language models.
Direct access to pre-digital history is drying up rapidly, why accelerate that for incremental benchmark gains in a domain that isn’t even relevant to the most useful forms of a nascent technology?
A physical book in a museum archive is useless for 99% of the worlds population even if they really wanted that specific book and were able to find it, as they'd have to arrange for access, then travel (at incredible expense) to access it.
Maybe they could ask the museum to digitize it, but that's still going to be days of delay and tens of dollars of cost to access parts of that book, if the museum even offers that service. If we go slightly beyond your "melted into" statement, the chances of the book becoming useful to the public are much higher in the AI company's digital archive, which might turn into something like what Google Books could have been, given the right incentives and copyright law changes.
And of course that presumes that the TV guide is going to stay in the museum rather than been thrown out as part of curation (or realistically, long before it makes it into a museum). Neither museums nor archives hoard everything, throwing stuff out is - as far as I know - one of the key jobs of an archivist. And a 1994 TV guide, while useful to understand what it was like to be alive in 1994, likely doesn't contain much unique information. You don't need that specific guide.
If there are 52 weekly editions, of 10 different guides, you would likely get most of what you want from any one of them. And for the parts that you wouldn't - there's a good chance that you'll have a much easier time getting the essence of this knowledge from the anonymized data pool that all the content was melted into, rather than chasing 10 different museums to find the original magazines.
https://genome.ch.bbc.co.uk/about
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
There is also a legal element. If they kept the physical copy around after scanning the argument is that they're making copies of the book which puts them on tricky legal ground. By destroying the physical copy they can argue that there is only one version of the book that now exists solely in digital form, so this usage is better protected under fair use.
I’m not aiming this at you directly by: ISBNs or STFU
Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this leads me to believe it’s people worked up about the idea of destroying (except it’s not destroying, it’s transforming, a fact often ignored) books, books that it’s not clear at all there is any strong demand for.
People want to invoke things like F451 but it doesn’t compare in the slightest. It’s like when people get mad about libraries throwing away or otherwise liquidating books that no one is reading in order to bring in books people want to read. People get all up in arms about that as if a book itself, in isolation, is inherently valuable or worth protecting. It’s not. If no one wants to read it then what value does it have? The impetus is on the people that think the book has value, it’s on them to carry the torch, to preserve what they think is worthy.
It would be like a company going to a yard sale and buying unsold/unwanted items to 3D scan them and destroy them in the process. This isn’t breaking into the Louvre and destroying one-of-a-kind artwork.
BBC good enough for you?
https://www.bbc.com/news/articles/cp3rprx2wl4o
"A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
"But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
"It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them. This isn’t complicated. Amazon/etc aren’t breaking into museums and libraries, they are buying books on the open market.
If these books are so rare and important, then why has no one cared until now to actually preserve them?
You're repeating the seller's contextual point, as if it's a counterargument. Do I need to explain to you that what makes the lone surviving 18th century edition important is that there aren't 75 copies of it in museums?
> And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them.
It's good to see you agree any digitization of this category of book should be non-destructive.
> If these books are so rare and important, then why has no one cared until now to actually preserve them?
(Lastly for realz this time, eh?) Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.
> It's good to see you agree any digitization of this category of book should be non-destructive.
I don't. Digital or physical, it's the same, there is no difference in my mind. I have many paper books but they are art, not functional, I also have the ebooks which is what I actually read. The _ideas_ are what's important, not a dusty, decaying shell in which the ideas are contained. I wouldn't shed a tear over every library digitizing their books and destroying the physical versions, nothing is lost. More importantly, once you've bought something it's yours, yours to read, yours to display, yours to destroy. On hacker news, of all places, the people fighting _against_ first sale doctrine is appalling.
> Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.
Our definitions probably differ here but preserving is not storing a book, preserving is ensuring that even if this copy is destroyed the ideas inside live on. I think that people that hoard (actually) rare books without a thought or care to making sure the text inside is preserved for future generations out of some desire to simply own something rare are the actually monsters here. And let's dispose with the notion that booksellers are "preserving" books, they are holding inventory, inventory they were happy to sell to Amazon/etc. If AI companies were raiding museums at gunpoint we'd be having a different discussion. They are buying books for sale, if they keep them in a library at corporate or scan and destroy them it makes no difference.
Also, who’s forcing AI companies to “ingest” books in such a destructive way?
Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.
Anna's Archive, for one, would be more than happy to host at no cost to the author.
The law is currently forcing these companies to destroy the books after scanning them.
No, the law is stopping them from digitally sharing their scans. They are perfectly capable of reselling or donating or storing the books they buy. (Wasn’t Amazon originally a book seller?)
digitize: no, there's no burden to do this
freely distribute: no, there's no burden to do this
use it or lose it on the copyright: yes; if a work is copyright but a publisher doesn't want to make new copies because they won't make money on it then it should be free to copy
Nothing forces them to shred books, they do it because it's slightly cheaper that way.
They can contact the copyright holder and ask/buy a license to make multiple copies.
Books that are rare of have historic significance will surely be in museums or libraries and not going away for pennies.
On the other hand, there are grey zones here, the lord of the rings books (still copyrighted and easily obtained pretty much everywhere) have been translated into my language many decades ago, and many of us read and liked those translations, but when the movies came out, a new translator did a new translation, where they changed a lot of things, including the last names of bilbo and frodo (Bogataj->Bisagin) and the Shire (Grofija->Šajerska), and the old version is sadly available only in paper form on second hand markets. On one hand, copying that if you only want this specific version would not cause a lost sale, on the other, you can get new translations (or english originals) pretty much everywhere.
Why are AI companies forced to shred books?
The OP article sounds quite opposite though - that AI companies are doing exactly this - destroying books so only they have the scanned content.
Unfortunately 50 years after the death of the author (or 50 years after publication for corporate owned works) has been locked in as a minimum term through international treaties, so it'll be somewhat hard to lower it beyond that, but many countries (including the US) enforce much longer terms, so that would be a first lever that could be applied quickly.
Maybe countries could could also establish an exception for out of print books offered to the public for free, or a general "library exemption" for public archives after a certain number of years?
I'm sure one of the AI companies would be willing to host a LibGen style library as a PR measure if legally allowed (with sign up required for rate limiting and as an extra benefit for the company to get daily active users).
Biggest problems:
I work in this space. I now require that before scanning a video is made carefully flicking through every page of the item so it can be checked after scanning to ensure all the pages are present and in the original order.Even the big libraries fuck up. I wanted an intact copy of Harper's Weekly from 1900 that has a big fold-out map in it. It's not clear to the libraries scanning this issue that the map is missing from their copies. None of the copies for sale from dealers have the map. Even when it is still glued into the middle it gets missed by industrial scanners. Google's scan only includes the (blank) back of the folded map.
Luckily GPT was able to track down a copy in a university special collections and fired off an email asking them to scan it. I just got the scan today:
https://imgur.com/a/vgMkM7b
(preview size, they sent a 500MB TIFF)
Now I can reassemble the issue and upload it.
I spend a lot of tokens getting LLMs vision tools to find the missing pages in vintage items and then try to reassemble them from other scans where available.
I'm also splitting up volumes to reupload. A lot of periodicals are only available online as giant multi-gig volume PDFs with all the issues in one file. I have a separate app I wrote to scan all the pages looking for covers so they can be split into PDFs and then identifying the volume/issue/month/year data from the cover or title page.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
I support Anna's Archive, by the way. Information wants to be free.
https://annas-archive.gl/donate
To watch the world cup I had to spin up a VM in Brazil to watch it with Portuguese narration because the free transmissions are region locked.
I would gladly pay 5 bucks for it if it was possible otherwise and avoid the hassle.
On Brazil the world cup was being transmitted on youtube. in NL only on traditional TV channels or Online for the same channels (all for free but in Dutch).
And literally as I write this I receive an email saying that my youtube premium was raised from 33 to 38 EURO. So there we have piracy getting juicier and juicier.
My concern is that less popular content is just erases, lost in mergers or lost in massive datacenters, never to be seen again.
I can already think of a couple examples I've run into in the audiobook world. I have copies of Douglas Adams himself narrating his Hitchhiker's Guide to the Galaxy and subsequent books. As far as I know, these recordings are not available for purchase anywhere. I got them because someone was kind enough to upload them. I would have preferred to buy them but didn't have that choice.
Some of Iain Bank's audiobooks were only available in Europe for a time (and maybe still are). For those I was able to convince Amazon I was buying in Germany at least.
Access to _all_ art and "culture"? For free?
But that's not what many people defending piracy are doing. Many of them do indeed have great libraries nearby, and art and culture on demand at prices they happily spend having food delivered to their house.
Given that they have to destroy them anyway, they're obviously also going to use the much cheaper destructive scanning (cutting off the spine and using a feed scanner rather than carefully turning page by page).
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
I imagine since the law recently cost one of them truckloads of money for their violations of it?
There is nothing that says you have to destroy something because you scanned it. This argument has been confusing me since I've seen this pop up.
It's also probably more convenient for them to destroy the books as opposed to trying to find space to store them. Knowing those companies, most likely they'd be just stuffed into some warehouse to rot after a couple years.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
Like, the current thing in the news cycle is a poll that young people are now more worried than hopeful about AI. Which sounds scary, but my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..
There are were polls about people being "more worried than hopeful" about satanic cults or alien abductions? With the youth being the most worried about satanic cults? Just like that knee-jerk "it's the bigger bubble", that makes zero sense.
As per the GDC 2026 State of the Industry poll, "52% said gen AI is bad for the industry, nearly double the 30% who held that view last year". But sure, everybody but HN, LinkedIn, and X bros are just luddites clutching pearls in their tiny bubble. They're the weak and stupid ones, and that is why the stupid shit said about them, day in and day out, isn't actually stupid. It all checks out.
What if people like food more than AI? Have you considered that?
The crux of the issue is that it is mostly a PR problem. AI is amazing but being promoted, in the eyes of many, by the worst people imaginable. Very akin, and overlapping in many ways to the crypto crowd.
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
There's no "2022 split", and the "untouched by machines" bit came from the marketing blurb of a company offering book scanning services - not the AI labs themselves.
At the AI lab level: the book scanning seems to be driven by copyright concerns, not data contamination concerns. There was a concern about AI contamination, but there's no measurable performance loss from ingesting post-2022 data with minimal filtration, and some tests attribute small but persistent performance gains to post-2022 AI contamination. It's unclear where exactly do those gains come from.
Why is all high quality text useful? The "inverse problem" framing is that all text reflects the thinking behind it, somewhat, and by learning to reproduce it, LLMs implicitly learn to reproduce some of the thought process too. They don't just memorize the dry factual knowledge, but also learn how that knowledge fits together, and how to reason about that knowledge - both in the specific case and in general. And that "in general" then surfaces in an LLM's ability to generalize. Which is very desirable.
So with this one copy BS are you not allowed to have backups of the data?
No data => No models => No competition.
> ChatGPT […] originally released on November 30, 2022
https://en.wikipedia.org/wiki/ChatGPT
> On March 24, 2020, following shutdowns caused by the COVID-19 pandemic, the Internet Archive opened the National Emergency Library, removing the waitlists used in Open Library and expanding access to these books for all readers. More than one user could borrow a book at the same time. Two months later, on June 1, the National Emergency Library (NEL) was met with a lawsuit from four book publishers. Two weeks after that, on June 16, the Internet Archive closed the NEL, and the prior Open Library CDL system resumed after the 12 weeks of NEL usage.
https://en.wikipedia.org/wiki/Hachette_v._Internet_Archive
-- Thos. Jefferson
A partial solution there is of course a larger budget and a "last copy" policy where the last copy of a text at least is stored away in deep storage against a future loan.
The local libraries also accept book donations for an annual fund-raising sale.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
https://www.techbrew.com/stories/2026/01/28/anthropic-ai-boo...
https://www.courtlistener.com/docket/69058235/554/21/bartz-v...