> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.
Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
Yeah, which is completely fine. There's a major difference between:
A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:
- This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.
- The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.
- The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?
B: Company uses freely available scanned copy of the text:
None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.
Then give up training on antiquated books. Why does an LLM aimed at providing utility for people living in 2026 need to be trained on rare (thus probably obscure) texts of yore in the first place?
Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity.
We've been seeing that headline for a few weeks now and I really don't understand the problem.
Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.
Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.
So what's the problem here exactly?
Also from the article:
> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
> A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies.
If you've read the articles covering this issue, you'll be aware the concern is over the fate of rare and out of print books, rather than your straw man (ie those available in 'thousands to millions' of copies).
> "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
> "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
> "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
Is there any indication that Anthropic is destroying books from the 18th century? Even the BBC quote is a conditional. Emphasis added:
> "It would be a much more significant problem IF one like that were to be bought for destruction, having survived this long."
I agree with the sentiment of course but it is really a huge IF they are doing that.
IF they wanted to train on, say, Leviathan by Thomas Hobbes, why buy an expensive edition from the 1600s when they would get the same text from a Penguin edition for a fraction of the price? It gets much cheaper secondhand too of course.
I'll go further, why would they want to train on expensive rare and out of print books? Are they, perhaps, competing on an AI benchmark based on extensive medieval knowledge of the cosmos? There's been a lot of pearl-clutching about lost obscure knowledge but y'all really reckon that kind of knowledge is valuable to LLMs?
It is also legal to buy potatoes and burn them, nobody would care if I do it. But if I buy up a food supply enough to feed a country and burn it it would be wrong.
> We've been seeing that headline for a few weeks now and I really don't understand the problem.
It’s powerful symbolism. It reminds me of that tone-deaf iPad ad that sparked outrage in 2024. The one where all the cultural artifacts were crushed in an industrial press to make a soulless slab of glass. And that was Apple, who is generally well-liked by the public.
> I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them.
Of course you can. Nobody is saying these companies aren't allowed to do what they're doing.
But what they're doing is disgusting and something I can't forgive. It's an escalation of the attacks against society that these companies have been engaging in from the beginning. This isn't about legality, this is about what's right.
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
> The main question is why aren't they leaking it to AA themselves?
Why on earth would they? Ultimate point for these companies is to make a ton of money, obviously they won't shoot themselves in the foot and give away whatever advantage they have, especially not to a free archive which is about doing good in the world, which probably isn't profitable enough for a company to care about.
Exactly. They set up this operation specifically to comply with the letter of copyright law and defend against publisher law suits. “Leaking” to Anna’s Archive is the last thing they’re going to do.
> The main question is why aren't they leaking it to AA themselves?
How is this a question at all? They’re scanning books because the courts determined that it’s the only way to use that data. They are forbidden from using digital copies found on places like Anna’s Archive. They must acquire and scan the book.
They cannot redistribute the book. The Internet Archive tried that and the courts shut it down. You cannot scan a book and share it without violating copyright law.
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collectively. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of that knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
I'm only one person, but I scan old books that had an impact on me growing up, and upload them to archive.org. Thankfully there are others that do the same. (And to be sure, FWIW, these are books that have not been printed for about 50 years—I suppose the software community would call them abandonware.)
If they're 50 years old they're young, and archive.org will likely block access. If they're not already on annas-archive (or the copy there is trash), your best bet is an anon upload to libgen.
They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying
> The print original was destroyed. One replaced the other.
So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.
Yes but let's continue using Claude to write code because we suck at programming. Really the only way out of this is to STOP NOW using AI and use our brain instead. These company will just shut down if we stop using, and thus paying, for their services.
Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).
I hate to be the bearer of bad news but you really do have to assume the worst about any of these "AI" companies, especially the large ones like ChatGPT and Anthropic.
They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.
A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.
Are you really advocating for assuming things with no evidence, by presenting no evidence for why one should do so? That’s not especially rigorous thinking.
>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.
Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.
no matter how you look at it, this is a systemic failure. if as a society we're going to mass scan our history then we should be building an archive for the future. not using availability of information as a moat. not doing it over and over again and throwing it away because of some odd rules to protect someones market position. not using it as an excuse to put paywalls around 80 year old field guides to field rodents in western massachusetts. not taking texts that had limited value and mining them for turns of phrase to be piled up into a useless grey goo.
>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
How tho?
Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take them home, so the seller doesn't have to pack them for the trip back. Just the other day, there was a whole bin of books in front of a shop, offering them for 50 cents a piece. They will be destroyed anyways.
Unless they are buying and destroying really old, rare books or important small-print books, it is not much damage. It is not like they will buy "all copies of all of the books", just one. And its just that the data in physical print most likely hasn't been used for training, so this can help you find more unmined quality data. Nobody is stealing your books, preventing you from buying more or destroying all copies of a single book.
And some of these books would rot out of circulation or be destroyed anyways. Some people throw away 80-100 year old books on the regular, as they might just be unimportant to them or the world in general. And once the last copy is thrown or rots, that book will die forever. This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.
Good to see someone making this point. I'm confused by the panic, because they are making it out like AI companies are destroying every copy of the book. They only need one, and they destroy it after scanning it only because they don't want to store them all. And storing or archiving all these books is not a trivial task.
I didn't fully understand why but apparently there is also a legal reason to destroy the books, it makes it them less likely to be considered copyright infringement.
Not really. Selling the book onward does seem legally dubious but legally nothing (yet) prevents you from storing the book in a warehouse. obviously it’s cheaper to dispose of them.
I think under first sale doctrine you have a much stronger case with destructive scanning. Google Books, HathiTrust, and Internet Archive's book scanning project have had a lot of legal expenses.
Yep, this feels pretty much it. Looking at the "rare books", it was books that nobody would care about or would just rot away anyway.
Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.
I think one of the major things you learn as you get older is that there is a huge abundance of people who say the right thing, and a much smaller group of people who do the right thing.
The internet made this even worse by celebrating people who only have to say the right thing.
The increase in people who believe that perception is reality has the logical consequences of people shouting their opinions loudly enough so that it becomes 'the right thing'
Knowing and enacting are two distinct things, "you should do what I want you to do" means you're already doing the wrong thing, and yes, there is an agreement on what the right thing is, its what is ethical, that is, what non profits and archivists are forced to do.
I can’t believe that in the multipolar world of 2026, where hundreds of conflicting world views coexist, where universalism is being disproven on a daily basis, it’s still possible to read things like “there is agreement on what the right thing is”.
This is as blatantly false as claiming that the Earth is flat, and the fact that there is no such agreement (descriptive moral relativism) has been firmly established in philosophy for well over a century.
If you think ethical arguments like "murder is bad" or "human knowledge should be preserved" and a demonstrable falsity like "the earth is flat" are equivalent arguments then you are so far gone you might as well be a flat earther.
Blind futurists and AI cheerleaders scare me on how cavalier and how many crimes against humanity they ignore.
Just go to a local library and ask a librarian or anyone who has a lot of books and tries to give them away; sadly, in most cases, they pick the valuable ones, and the rest just get sent for destruction (Burning).
The difference everyone is missing, is that now there is commercial incentive to burn books aka destroy them after scanning. It's now a profitable business to do so, not something that only has to be done to clear up space
Weeding (deselection works) is a fundamental part of collections management. Every trained librarian is going to understand this.
The interlibrary loan system has mechanisms in place to make sure the member libraries keep two copies of each work in each region. Collection managers consult these databases during weeding to make sure they don't deaccession the last copy.
If the monograph was never collected by a library and it gets caught up in a destructive scanning project then I guess it was pretty "rare" in a literal sense. "Rare Book" in library land is sort of a term of art and I'm not sure if the books in these destructive scanning projects meet the criteria.
I'm not a bsky person so I didn't click though but library books aren't "retired" to a farm upstate. My SO works at a library, they are mostly shredded. This is a nothingburger
> The example book of Old books of agriculture is probably not that important today
If I may be flippant, not to you but to the sentiment, skill issue.
We're about to enter an era of climate instability that's going to cause wild fluctuations in the ability to grow food across the globe. Historical agriculture data AND data about confounds is crucial for figuring out what strains outside of our current mostly mono-strain agricultural supply chain could be cultivated.
And that's just one use case out of thousands; what if you want to understand and reconstruct technology adoption from that era? What if... you just want to learn what your ancestor was doing at such and such time? What if you want to find clever techniques for robot arms to work with food crops in space?
> Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
I don't understand what you're trying to say here.
"No-kill shelter" are very rare; most still have to kill by law (overpopulation), and they instead tend to not take in more than they can handle. They are very controversial because they are more of a "moral" than a "practical" solution to animal overpopulation. It is written like it is the "end of the world", just like your random rambling about "climate instability" is just the natural course Earth has been on for its existence, with melting ice.
Random article about powder and stone, completely irrelevant.
The old book mentioned was just a farmer's book; you can find all this history in the area's historical records.
"Skill issue" for one is that it's used wrong, but it tells me more that you tend to spend way too much time online and from your insight, very little knowledge about the work of shelters and Library.
Because destroying the book makes it possibly inaccessible permanently because we have no idea how long anthropic plans on storing the digital copy, if at all. A lot of people assume they would for future training, but you don't know that.
If they at least didn't destroy the book, someone could purchase it when anthropic eventually goes belly up. Hopefully someone will at least be able to purchase their digital scan and hopefully the scan is of decent quality and clearly indicates the provenance of the text.
> we have no idea how long anthropic plans on storing the digital copy, if at all. A lot of people assume they would for future training, but you don't know that.
Books are considered super high quality training data. Anthropic has no reason to get rid of this data that 1. They’ve spent a ton of money on and 2. Will remain useful indefinitely for training LLMs.
> If they at least didn't destroy the book, someone could purchase it when anthropic eventually goes belly up.
The same applies for a digital scan? The information isn’t any more likely to be lost.
Yesterday you could have purchased one of these books and read it. Good luck doing so today.
I don't know why you're so quick to assume this will all just "work out" such that the scans are ultimately accessible. Arguably the most likely scenarios are either Anthropic survives and holds them away in perpetuity or Anthropic fails and they are sold off to the highest bidder who does the same.
A lot of them are rotting. We are not talking "one of the three living copies of the first edition of Joyce's Ulysses". Rather "1956 statistics of the cultive of yuca in 'some small village from Mexico': a boring analysis". Those books have value to train LLMs as they are 100% free of AI text, but has been collecting dust (or rotting) in someone's room for decades, and no human is buying them even for 10 cents.
Also, Anna's text implies that the books are scanned and then mischievously destroyed so nobody has access again to the content. That's not the case: the books are "destroyed" before scanning, by dissasembling them in pages so they can be feed to the scanner. Scanning while keeping the book intact is difficult, as you need to software-unwarp the page before OCR'ing it, and expensive as you either need specialized scanners or humans doing it.
> This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.
But that's just it, it won't live on forever because, from the perspective of preservation, training is a lossy, noninvertible transformation. The LLM cannot legally produce the book verbatim, it will only spit out a regurgitation of the information, chopped and mingled into a broad information space.
Furthermore, these "magic machines" are not the property of the public. They are owned by a handful of corporations who want to charge you continuously for every token output by the machine. So, not only is the original text locked away forever behind company walls, you now need to pay for access to an approximation of the original contents which you can no longer even verify as being correct because the source is no longer accessible.
If you are cool with this, from a cost perspective you are cool with a deal whereby I trade you access to a definite resource for a one time fee of $N for, instead, a perpetual cost of $M to you every month/day/hour for access to an amalgam in which you cannot even determine what proportion of the resource you are actually getting. You're basically saying you're cool with me selling you some unknown portion of wine for a monthly subscription price instead of selling you a definitive amount of wine for a one time fee. lol.
Are you under the impression they scan the book, train on it, then destroy the digital copy? Because that's not what's happening. They scan it, and hold it forever to train future models on. The scan still exists, not available to the general public but that's no different than if they had bought the books and kept them in a private library closed to the public.
>it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge
This sentence reminded me of “The Things” by Peter Watts. The Thing in that short story believes it’s actually the good guy and decides to commit “violent integration” for the sake of humanity.
I’m not sure I buy the links premise that Anthropic is doing this with malice, but if it were, would be an eerie parallel.
A small change in the copyright law would fix this problem. Something like:
If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.
Wait, so the library of congress is suddenly responsible for probably petabytes a day of incoming scans? To what end? Do they have to index it and make it available? Do they have to check the accuracy and integrity of the scans?
How does this help anything, except create more work to throw in the trash?
That revolves around print on demand for books that are out of copyright or where the copyright has been abandoned.
>
Background
Valancourt Books is a print-on-demand independent publishing house specializing in rare and out-of-print books. Valancourt had not registered its books for copyright as the Library of Congress already had original-edition copies of the books Valancourt republishes and any new material in its publications was limited to notes and introductions.
If you want a physical copy of The Sorrows of Satan, you can buy it from them.
> The Copyright Office has stated that it would modify the language of its deposit demand letters and withdraw its demand for copies if the Copyright Office was notified of the copyright's abandonment.
> Several legislative changes have been proposed to address all elements of the case: changes to Section 407 to tie some legal benefit to the deposit, monetary compensation to copyright holders for depositing books, and regulation for a simple and costless method of copyright abandonment.
That doesn't change that if you were to publish a book today (or for that matter, have published a book in the past 100 years in the US), you are required to deposit a copy of the book with the Library of Congress.
Actually - most jurisdictions that issue a publisher a unique root ISBN number have a stipulation that anything new published using that number must have a copy sent to them.
Looked into this a decade ago for publishing eBooks via my personal corp when eReaders and ePub were starting to hit big in the mainstream.
Puts the burden on government to store what is probably 90% worthless material.
Copyright should really be amended so that once out of print and a grace period it’s free use. I am probably more of an anarchist in this regard. Similar to my belief that anyone should be able to ingest any data you put online, once a book is no longer being print it should be able to be used for commercial or personal use for free. Similar to a generic drugs.
There is far too much garbage that gets published, let the collective hive mind figure out what is valuable.
Since we are talking about a US perspective do you have evidence that backs this up? It just comes across as an empty statement. The government is the will of the people and I personally like the idea of fixing copyright instead of making the government store how to use windows 95 books.
Sure some governments and opinions would say so but you’re making a statement of zero impact. Fix the underlying copyright laws don’t create more rules.
> All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17).
> This law requires two copies of each work published in the United States be deposited with the Copyright Office within three months of publication. Works deposited under this law are for the use of the Library of Congress. Usually, deposited copies must be the “best edition” of the work, which means they must conform to the Library of Congress’s preferred specifications.
> Mandatory deposit applies to any work published in the United States. This requirement does not apply to works first published in a foreign country until they are published in the United States. Copyright registration is optional, but it provides additional legal benefits and fulfills the mandatory deposit requirement with the submission of the required copies.
You don't need a massive government program for this.
Just post on r/DataHoarder: "Free 16TB NVMe SSD to anyone who indexes and mirrors the entire out-of-print 20th-century physical archive." The problem would be solved by next Tuesday. With probably 10x redundancy and people willing to do it for free for fun.
Wrong end of the pipeline; we should instead demand digital copies of media be sent to the Library of Congress in order to obtain copyright, along with a registration fee to pay for indefinite storage and other costs. Registration should be mandatory if you want copyright. For things like books where a machine readable text format existed, it should be mandatory to include (so no requiring OCR). Access to the archive should be available for research use (including ML training) at cost.
> All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17).
> This law requires two copies of each work published in the United States be deposited with the Copyright Office within three months of publication. Works deposited under this law are for the use of the Library of Congress. Usually, deposited copies must be the “best edition” of the work, which means they must conform to the Library of Congress’s preferred specifications.
----
> Acceptable Formats for Deposit of Electronic Works
> The deposit of electronic works is arranged with the Acquisitions & Deposits division.
> For electronic-only works, submit the best edition in accordance with the formats listed in the “Electronic-Only Works Published in the United States and Available Only Online” section of the Best Edition Statement (PDF, 135 KB).
> For works subject to a grant of special relief, unless otherwise specified, the Library will accept an appropriate “preferred” format listed on the Library of Congress Recommended Formats Statement. Such files must contain no measures (such as digital rights management [DRM] technologies or encryption) that control access to or prevent use of the digital work.
> For more information about electronic deposit, see the above FAQ “When can I make an electronic deposit of a work?”
---
> When can I make an electronic deposit of a work?
> Works may be deposited in a physical format in accordance with the Best Edition Statement, which can be found in Best Edition of Published Copyrighted Works for the Collections of the Library of Congress (Circular 7B) (PDF, 135 KB).
> Works may be deposited electronically in certain circumstances:
> The Copyright Office issues a written demand for an electronic-only book or serial. If your work is published only online and the Office sends you a written demand for mandatory deposit of the work, you must deposit the work electronically.
> The Copyright Office offers you electronic deposit as an alternative to depositing a physical copy of the work. If you receive a letter offering special relief to deposit a work in an electronic format instead of sending physical copies, follow the instructions in the letter or agreement.
If the Library of Congress has a digital copy, it would be easier for them to distribute the work after the copyright of the work expires. That would be a public benefit.
> it would be easier for them to distribute the work after the copyright of the work expires
Copyright does not expire for a very long time. Harry Potter and the Sorcerer's Stone was released ~30 years ago in 1997. It remains protected for the duration of the life of the author (J.K. Rowling) plus 70 years.
Given actuarial tables from the UK[1], this works out to be around ~95 years from now (~2120).
Certainly. But the rare books under discussion are closer to the end of their life and less likely to have been already digitized.
The library of congress does distribute some digitized works that are out of copyright. And it does digitize some works for archival and distribution, but having additional works digitized for (eventual) public use could be nice.
> But the rare books under discussion are closer to the end of their life and less likely to have been already digitized.
It is not at all clear that this is true.
The number of books published every year is growing rapidly. According to Bowker the number of books published every year has increased ~15x in the past two decades [1].
Because of this, I suspect that the median age of these books we are discussing is below 30 years.
In theory the Library of Congress could "lend out" digital copies like some library systems do. This would be especially helpful for rare books since it is less likely multiple people would want the same book concurrently.
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria
Let's view it realistically here: AI companies are parasites. Them destroying books to dumb down mankind, absolutely fits into the destruction of the library of Alexandria.
Having said that, I think the day of physical hardcopy of books, is not necessarily over, but will be heavily complemented via digital storage. For instance I only keep books that I may re-read later or read many more times, e. g. thick science books. Many other books I can keep as .pdf file without a problem.
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria,
These comparisons are starting to get ridiculous. Why are so many people assuming there is exactly one copy of all of these important books available, and that Anthropic is destroying the lone copy?
Your local library throws out books every year and nobody thought twice about it.
>Why are so many people assuming there is exactly one copy of all of these important books available, and that Anthropic is destroying the lone copy?
Why are you assuming that each book gets scanned exactly one time and then never again? And why are you assuming that out-of-print books remain easily accessible so long as not every copy has been destroyed?
>Your local library throws out books every year and nobody thought twice about it.
When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.
As someone who has gone to many many used book sales over decades… many of the books at a sale never get sold, guess what they usually get tossed in the dump. This includes your local library book sales. Books are heavy and worthless. Cheaper to throw away the ones that nobody picks up in a sale.
I know it comes at a shock but truly most books are absolutely worthless.
> When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.
Many books aren't lent and not bought and most libraries have limited space to store such books, thus they go where old paper goes.
Of course some rarely lent books are important and for the one person asking for it in ten years really valuable, but many still have to go.
No. Almost none of the books you donate to the library get to the shelves. If they can sell it, they will. Otherwise they are thrown away.
I even read on a web site of a librarian that their library had stopped accepting donations because "patrons should know how to throw away their own trash".
> "patrons should know how to throw away their own trash"
That makes me a lot sadder than AI companies destroying the books but at least preserving (most of) their knowledge.
They are ridiculous because the reality of the situation is boring. Boring doesn’t drive engagement. So, all the clickbait headlines and ragebait comments imply scandal. If they didn’t, they wouldn’t get attention.
After being clickbaited and ragebaited, media consumers feel deeply anxious and angry. But, explaining that they are angry over a boring situation feels silly, not righteous. So, they give summaries, impressions, sometimes extrapolation of the bait they have been consuming. That feels righteous.
This observation applies to a wide variety of topics trending in the various media every day. Distinguishing injustice from ragebait unfortunately requires non-trivial effort from the reader.
Does china developing ai powered missiles sound ridiculous? It sounds ridiculous to me because the US is surely doing the same, It would be boring if we knew the us was doing the same. It is emotional to think about china destroying the US when it is just the antithesis of American exceptionalism. It reminds me of what Dario said about their stance on open source.
That doesn't sound ridiculous at all (maybe reductive but not ridiculous). And the US doing the same doesn't make it any less interesting either. In fact the US's use of AI in developing target banks in a recent conflict caused quite a lot of consternation when publicly revealed. Ukraine has also recently used AI enabled autonomous drones to make an entire area one big "kill zone", no humans required. Not so ridiculous or boring if you ask me.
They're not destroying rare manuscripts or incunables.
They're destroying one (1) copy of a mass-produced item for each AI company.
Public libraries destroy millions more yearly as a matter of routine.
This is just part of a CCP-aligned moral panic, along with the water use nonsense, and similar with the soviet-aligned moral panic that destroyed the civil nuclear industry 40 years ago.
I was with you until the water use. You're misinformed. There were at one point at least several data centers set to use evaporative cooling on well water.
Notably since all the controversy many data centers are very loud about being closed loop and with significant consideration given to other local impacts as well.
It's possible for a datacenter to use scarce well water irresponsibly.
They don't have to, and the vast majority don't.
The lie and moral panic is that all datacenters necessarily waste precious drinking water; it's patently false and used by agitators to push, unwittingly or not, a Chinese Communist Party agenda.
is it actually? it feels more like US propaganda that we have to let our oligarchs run roughshod over us because of what we imagine the big bad CCP might want.
i dont think the CCP cares whether there's data centers in rural america.
Regulation that requires closed loop cooling seems simple enough, same with lots of the other problems people have with data centers:
* sound and infrasound under x DB
* no air quality change
* must pay to build out electrical infrastructure
etc
its not to the CCPs benefit or loss to make sure the data centers are built well if they get built
No, that’s the hyperbolic reaction clickbait wants from you.
Not all rare books are valuable. Someone’s self-published junk sitting in the garage is NOT analogous to the library of Alexandria.
Many, most, maybe all of these “rare” books are being scanned instead of just being recycled.
Not a big Reddit fan but there was a great post there from someone in the book industry talking about how non-industry people often give this great moral weight to ever book in a way that is totally disconnected from reality.
All I’ve read, as far as sources go, is a number of rare book sellers saying they’ve had a big uptick in huge orders with no price haggling. Apparently that’s peculiar. And some of them seemed a little concerned.
Now I’m certain they’re not chopping up Davincis notebooks, but I’m not certain there aren’t some that would make people wince.
And I don’t have any reason to think some reddit librarian knows what’s going on, if anything, either way.
I heard from the first stories that virtually all of these books being ordered have ISBN numbers. Books that are rare that have ISBN numbers are rare because no one wanted them 99.9% of the time. Somebody wants every book, but you'd spend many, many years finding that somebody.
A flagship LLM today is trained on tens of trillions of tokens, the equivalent of hundreds of millions of 100k word books. No human has that kind of appetite.
As an aside, the entire Google Books corpus is generally estimated at tens of millions of books.
Exactly. It’s the funniest thing. These books are only being bought by AI companies. It seems pretty clear they only have value to AI companies. Otherwise all these people bemoaning the loss of rare books would be… buying them.
“How dare these companies buy rare and valuable books that nobody else values enough to buy” is a self-canceling argument.
If you tell people who need digital text from books that they need to destroy books after scanning them, they're going to use destructive scanning and destroy the books.
Depends on the point being made about “historical precedent” and the lessons to be drawn from such.
Also, helps to clarify what exactly the commenter was referring to and possibly help distinguish the centuries-spanning decline of the Library of Alexandria from the violent fate of the Serapeum.
It's a fair point. I honed in on "the burning of" (original comment) versus more generally thinking in terms of "the loss of", because parent context here is "AI companies destroy…".
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
How can anyone say this with a straight face. The knowledge is not destroyed, it is transformed. You can make use of it today in the form of LLMs and the scans still exist. Nothing was lost. It's literally no different from them buying books and stocking them in a private library not open to the public. It's not called the Scanning of Alexandria because if it was, it wouldn't have made a blip in the history, Alexandria's libraries were burned, those books, that knowledge was destroyed. Then only thing being destroyed here is physical copy (again for the people in the back: a copy).
> Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.
Those same copyright restrictions are exactly what would prevent them from sharing the archives. Your beef is with copyright, not the AI companies who are (in this one, rare, instance) following copyright laws/rules.
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria
Hyperbole much?
Does the fact that they're being converted to an immutable digital permanent record for all time mean anything to you? Because as far as I know, the works lost to the Library of Alexandria were wiped out of existence, not simply transformed into a more durable form!
The idea that these corporations or any of the literal sociopaths that work for them give the slightest bit of a shit about "benefitting humanity" is hilariously naive. The one and only thing these entities care about is money, and making as much of it as they can. If they could get away with it, they'd commit every crime that exists if it meant they get a quarter of a percentage increase in their quarterly earning reports.
Are they? Or are they just using it for training and not keeping the digital copy afterwards? And even if they are keeping a digital copy, does that actually matter if they never release it?
You’re making a distinction here, but training is something you repeat for every new point release, so you need to keep the data if you want to use it for training.
Agree. While it certainly isn't the most environmentally friendly to render huge stacks of paper into waste, the real issue is copyright creating scarcity (inability to copy the thing).
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
I dunno, that "at least" worries me, digital actually more easier to be lost if it is under copyright, they just got deleted if they can not produce enough money. Physical may have higher chance to survive.
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
That's what state libraries are for, although I understand that sometimes is hard to wrap around the concept of using public money for something different than producing money.
So these books are not worth anything because they have so few copies but they are still worth including in only their models? Seems a bit contradictory
Not really. A trivial example: smut novels. I'm sure AI companies want them for training so their models work better as AI girlfriends/boyfriends, but I doubt much would be lost if the bottom 50% (by readership) of such books went into a woodchipper.
I ask "more clean air", you answer "why, did you deserve it, having clean air is not free, there's a real cost"
Somebody else asks "We need more accessible energy", you answer "why we bother with your needs, you're not efficient, energy belongs to more efficient purposes, you can live without that much energy".
I can continue with more, but hope you got another viewpoint.
Btw, I'm disgusted there are "humans" like you in existence. We definitely don't share the same cultural ancestry, and I hope ours will prevail at the end, rather than cold blooded, mechanical "brains" like those of your kind.
> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable.
So old books are valuable because old paper is valuable? I understand why the Gutenberg bible might be valuable but do we really need thousands of mass paperback novels?
You're confusing the worth of the book and its content. A book can be valuable (ie a rare bible print), whereas its content is not (we have all the bible variants copied).
>If they are truly rare, then they are likely not valuable,
Sometimes I don't even know how to respond to comments here. I don't want to be rude, but you just have to give this a moment of thought. Is all the media that you find valuable common? I know that's not the case for me based on my own experience.
Massively missing the point.
Having paper books isn't the point.
Preserving copies of books is the point, so history isn't lost.
Often those books only exist as paper copies due to their age.
I don't really care what happens to the paper copies, only that their content is preserved in a way that is accessible.
Anthropic's private servers aren't it.
Nobody's said it in this thread so I'll drop it here where you ask "Why save physical books?"—
The problem isn't with morals or copyright. What we're up in arms about is case law. Past rulings have implied that destruction of books significantly contributes to the process being "transformative". This encourages companies to destroy the books. I think this is really dumb.
Why save physical books? It's because the reason to destroy them isn't good. If you think there's too many bad books out there, that's a different argument. Maybe your fight is against consumerism, I don't know.
What you personally find important is not what everyone else finds important or inspiring. Destroying something takes it away from every single future human being.
They may have been made scarce by many methods, perhaps a low print run, burning in a revolution to suppress dissenting views, scanning to avoid competitors to get the same content and prevent other LLMs to reference, search or train on it, among others.
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
It is not a big deal. Since the invention of the printing press any important
book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
Read more about this. Depends on the definition of the "big deal" but from what I can understand the problem is that they buy rare things - which exist in just several copies - and they tend to buy _all_ copies.
This is a classic mistake. We have no way of estimating the future value of a given book. It's perceived current value (a large part of which is simply obscurity) may be low. But it's future value - to historians, ethnographers, to researchers seeking a specific fact or example of language use or a hundred other things - is literally inestimable.
To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved.
Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.
>Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.
I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.
OK... I'm going to assume good faith even though your wording makes it somewhat unlikely.
Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing ecosystem or the state of a particular political situation at a given time. Essentially all textual primary sources for events which are not currently considered important - but which we have no way of estimating the future importance of. One can continue to create countless counterfactual examples in this vein. My overall point is we cannot know what may be useful or even essential in the future, and knowledge should be preserved under the assumption that it is likely to be.
Historians frequently refer to this paradox - how everyday aspects of life are frequently not explicitly documented, since they're so obvious to the communities or communities of expertise that observe and carry them out. So it's actually incredibly important to preserve what seems like ephemera.
Hell we couldn't have AI training at all if we lacked the corpus of existing written literature - but there was no way any author could have anticipated this future utility more than a couple of decades ago.
1. If someone is acquiring books in bulk for bargain-bin prices and shredding them, they're books whose physical copies have essentially no market value and which, absent this buyer, were overwhelmingly headed for pulping or landfill anyway. Millions of books are destroyed every day.
Could one of these worthless looking books turn out to contain information historians care about in a 100 year? Sure. But that doesn't create an obligation for someone to pay to warehouse every extant copy forever. Physical Preservation has costs: space, cataloguing, handling, transportation etc. Archives and libraries have always had to make choices for this reason.
2. I'm not arguing that preservation has no value. The question is whether destroying a physical copy after digitizing it is a serious loss when talking about mass-produced printed material.
Was this the last surviving copy ? Is the information unavavilable in libraries, archives, other editions, scans, citations, contemporary works etc ? If not, nothing has been lost except one physical instance of a reproducible object.
Your film analogy worked a lot better because old film footage is often unique primary source material. A camera recording of a random brooklyn street in 1993 may literally be the only recording of those people, storefronts and circumstances. The nth printe dcopy of a technical manual is not analogous to that.
The medium obviously matters. One was designed to be cheap and mass produced. Almost anything published has copies in the thousands at least and the chances you are concerning yourself with the 'last copy' is minuscle. Film most certainly was not like that.
Assuming these texts exist in large quantities seems to definitionally contradict the term *rare* book. Plenty of books have small print runs or have reached out of print status.
One could obviously reproduce films too.
Ultimately, you don't know how big N is, and that's what matters. Part of the problem is there is no transparency around this. The population of all books != the population of a specific book or set of books. Assuming N is large when you have basically zero information on specifics is foolish. These books might be just as rare as a film for which only one source exists you don't know.
>Assuming these texts exist in large quantities seems to definitionally contradict the term rare book. Plenty of books have small print runs or have reached out of print status.
They're books that are no longer in print. Doesn't mean there aren't a lot of copies around.
>Ultimately, you don't know how big N is, and that's what matters. Part of the problem is there is no transparency around this. The population of all books != the population of a specific book or set of books. Assuming N is large when you have basically zero information on specifics is foolish. These books might be just as rare as a film for which only one source exists you don't know.
All of this is frankly irrelevant. Books that you buy in bulk at barging bin prices are books that have essentially no market value and were going to the pulp or landfill anyway. Nobody has an obligation to spend money to preserve every extant copy of every book forever. Millions of books are destroyed everyday. Even libraries and archives make these choices. I was simply pointing out the film analogy worked better, but it doesn't change anything.
Again, you're assuming a lot about this situation that we don't actually know to be true. For example, that these purchases are limited to "bargain bin" books that have essentially no market value. These may be cases we know about, that doesn't mean these are the only purchase happening. I would hypothesize that if we know about a few specific purchases there are probably many more we do not know about/haven't traced back to these companies.
You realize "market value" isn't the only kind of value too, right? Even books that are completely obsolete can have inherent historical value insofar as they enrich our understanding of the past and how human beings used to live. Contemporary market value isn't everything. There are plenty of works of scholarly interest that wouldn't command nearly requisite value to a NYT bestseller when you adjust for scale. That doesn't mean these books aren't worth keeping or aren't intellectually important.
Also, these books presumably even have some perceived value, otherwise these companies wouldn't think they were relevant for training, and wouldn't be spending money to purchase and scan them in the first place. The whole premise is that these books are valuable enough such that training on them will make the models valuable and make paying customers line up for tokens. This is an inherent imputation of value to these works on the part of these companies. You can't say, simultaneously "these books are worthless" yet somehow, at the same time, "they are worth something for producing artificial intelligence". Your assumptions don't make any sense.
> They're books that are no longer in print. Doesn't mean there aren't a lot of copies around.
Sure, but there also could be limited copies around. We don't know. We won't know unless these companies are more transparent about exactly what they are buying and scanning. It's that simple. If it's foolish to assume there might be limited copies, it's equally foolish to assume they have plentiful copies. We simply don't know.
> Nobody has an obligation to spend money to preserve every extant copy of every book forever.
I don't think anyone is claiming that. People are bothered because they (in my mind, rightly) recognize that old books are a historical record of human accomplishment and have some amount of cultural value. Letting a handful of companies further plunder the collective output of humanity and now, potentially keep it locked on their servers in perpetuity is something that should make you upset if you have any inkling of interest in the history of humanity or human accomplishment and if you have any sliver of curiosity about these things. If all you care about is the present and getting better LLMs, sure, I guess I understand why you don't care, but then I also think you are very unwise and myopic.
> Millions of books are destroyed everyday.
So what? Millions of crimes are committed every day, that doesn't mean we shouldn't be bothered by the crimes we hear about. Billions of pollutants are spewed into the environment every day. That doesn't mean we shouldn't try to combat pollution. That's a pathetic attempt at justification. Also, there are clear distinctions between processes like libraries removing works they can no longer store and companies scooping up books for training and personal gain.
It's not that individual companies buy all copies of a given book, but that there's more than one book scanning company, and they aren't sharing the scans with each other. The result: books that were rare but nevertheless easy to find for purchase (thanks to the internet) are now vanishing off of the market, becoming de facto no longer accessible to the public.
Good article, but I still feel unsatisfied because even it cannot find an example of a book that’s actually been lost because of the destructive scanning frenzy (it only lists books that hypothetically could be lost because there’s not many physical copies available for sale online.).
Since we don't know what was actually purchased and what was actually destroyed, how do you expect us to furnish an example? This would require the destroyers to admit it, and beyond that it would require all of them to admit it since more than one of them might have been responsible for the extinction of one text. Seeing as they were already keeping this operation under wraps, I don't see that happening. "possibly extinct because no copies available online" is probably the best we can do.
The distributed nature of the problem and the utter lack of transparency are huge factors here too.
These companies are buying from small book stores; surely these collector types would know if they lost any one-of-a-kinds? Or at least someone would be keeping track of extremely rare books disappearing (especially now since this matter has been public for weeks).
Regardless I feel like they gotta figure this out for optics reasons. “So and so books are lost forever to Anthropic’s servers” is much more outrageous than “Anthropic is destroying a bunch of books that have other copies” imo.
Why do you think that? Do you have evidence, or is this just a hunch? Why would Amazon waste money on uber rare books when there are thousands and thousands of not-so-rare books that could serve the exact same purpose?
Exactly, also all those physical books would anyways get molded, eaten by moths or just naturally decay. It's not like the AI companies are obliterating every copy of every single book.
> a digital copy with the ability of doing millions of copies is stored somewhere
somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”
the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.
> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.
archivists keep everything, because we don't know right now what will be important 100 years from now.
By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther.
Plus having the info part of a LLM makes it immediately available to literally billions.
I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.
>Plus having the info part of a LLM makes it immediately available to literally billions.
Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong.
If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions".
And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.
For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book.
The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly.
But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great.
That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.
> But, that little bit of data is a bit more data than existed before,
No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference.
> So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of.
That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy.
I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years.
> Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost
I understand what you mean but... "permanently lost" sounds dramatic. When I trow away old pictures, old drawings or pieces made by my son at school, they are lost as well. Not that I do that often, but it begs the question, should every 'ip' made by humans be preserved?
It sounds dramatic because it’s dramatic. You’re combining two things — whether something is, in fact, completely lost, and if it is something that should be kept. Something that should not be kept is still completely lost if it’s destroyed. It’s difficult to imagine they’d digitize it if it was worthless.
A small fraction of them is saved in the model. Far more is saved in the digitized copy as long as they keep it which they have plenty of incentives to do so (future training of newer models).
That's more than what happens if that book was burned or sent to a landfill, but less than if the book is giving a loving home.
>They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine.
My understanding is that this simply isn't legally allowed for these books. The original must be destroyed for the digital copy to not be copyright infringement.
>That’s a false dichotomy.
I pointed out there is a spread of possible outcomes and that different people are considering different outcomes and the comparison of if this is good or bad depends upon which outcome one considers. I even mention that both outcomes are sometimes right. That's about as far from a false dichotomy as I can see it.
> Plus having the info part of a LLM makes it immediately available to literally billions.
Isn't this a contradiction?
I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions".
> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed.
I don't if that is true: a lot of old books might still have copy or other rights associated to them, likely owned by author and/or publisher, directly or inherited, but often those who have the rights do not have digital or physical copies at hand anymore (some old books are, well, really old). Does Anthropic make sure to track down, contact and then share the digital copy they make with those who have rights on the work? If not, they are not making it in any way easier to re-print the books, while making their supply more scarce (they destroy existing embodiments).
> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed.
Or be useful and scan book for Anna archive and other shadow libraries.
Citizen's lobbying against megacorp is a mirage.
>But it's _closer_ to being widely available, not farther.
By what metric ? The copy is now guarded by a company instead of being on the second hand market.
You really drank all the Koolaid they had to offer.
Destroying the last copy of a book is OK because you shop for second hand books, and because some private company hold the last digital copy and have no incentive to make sure it survives. Damn..
> permanently locking human knowledge inside private corporate servers
History tells us that very few "permanent" situations are truly permanent.
Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.
I mean books get destroyed all the time, really they are a major pain in the ass to keep together, especially as they age. Paper loves to crumble. Insects think they are tasty. Floods and fire love destroying them too.
So physical books are rather non-permanent themselves.
I wonder if Anthropic could rent or sell access to their collection to the internet archive? It would probably be a good PR move (which they probably need right now), but I'm not sure what type of legal shenanigans they would need to do in order to not violate copyright law.
First of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans.
Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway.
Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why?
Why wonder? The answer is abundantly clear if you follow the news. If a book is copyrighted under U.S. law, scanning and destroying counts as a format conversion which qualifies it as fair use, so there is no need to negotiate with copyright holders. See Judge William Alsup’s decision. If Anthropic did not destroy the books after scanning it would have not won the lawsuit, and scanning would be illegal. If a book is already out of copyright then of course they do not have to destroy it afterwards.
Why would amazon buy "crap"? Surely they want their model to succeed and they want to train it on valuable input, no? They have more than enough resources to determine whether or not the books are worth buying. They have been in book selling for a long time.
I think you may be greatly underestimating the long tail. Several times a year I read sources that reference older books that I can't find online. When I am able to locate them, they often cost at least several hundred dollars, sometimes into the 10s of thousands.
Do you think that you are somehow special and unique in needing these books? If the books cost in the 10s of thousands, then there's obviously value to other people. I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus. All evidence I've seen is that they're scanning cheap books with no current value and no clear use to people today.
I have volunteered with a library, and probably threw hundreds of books over a couple week engagement from a university library into a shredder at the direction of a professional, academic librarian. Libraries are constantly culling books, the EXACT category books we're talking about here (old, never-read). This is happening at a much larger scale, so I would recommend railing against university librarians in addition to the AI juggernauts.
I'm just stating that not everyone is reading the most popular million books. And, there are many, many millions of books in the long tail.
A few years ago, after running into this issue several times, I looked into the economics of it to see if there was a opportunity to republish digitally. I found the sales data for some, the ones that had value were selling on average for around $50-150 dollars. Because the typefaces weren't modern, OCR wasn't scalable. Because only 40-50 were transacted each year, it wasn't worth the time.
It's interesting that libraries are purging them. Several times, I've found that the only available copies were at some random university rare book collection in middle of no where, and basically impossible to access unless one fly in - which isn't worth it. They are often donated by a benefactor and stuck there - which is why they're never read. It's not that the content isn't valuable.
What kind of books are they? In my case mostly historical documents by some relatively unimportant person who was highly important for a very, very niche subject.
They still contain valuable, irreplacable information. And, once the physical copies are destroyed, they'll probably be gone forever. An analogy: imagine you discover a really cool video game from 20 years ago. You love it and want to find the developer's previous work. But, you find out that the company was purchase by another company which was purchase by another, etc, etc. Sure, maybe the game still exists in some digital vault. But, more than likely it's gone because old things only seem have value these days if some influencer advertises it.
I think I get your point, and I’m sure there are ideas that will disappear forever.
The books we were disposing of were so totally inane that I couldn’t tell you what they were, but they weren’t taken out in like 30-40 years. Northeast US state R1 research university. Maybe we got rid of stuff that would come in handy!
Beware that the notion of "quality" is entirely different for AI companies: they don't seek entertainment, but sentences in a language to train an LLM.
I have books, but they are just objects. They're nice objects, but just objects.
Fetishizing books isn't going to help.
In fact, most of those "rare" books don't sell because nobody wants them. The AI companies are making them even more rare, so the booksellers should be thankful.
Honestly, it’s easier to find a good movie than a good book, because books are way cheaper to publish. These days, the quality of pretty much all kinds of content has become a problem...
> any important book has been duplicated by thousands, tens of thousands or even million of units.
Books in the former USSR display their print runs on the last page. "Important book" is a vague and arbitrary term, but rhere are works in whole fields (e.g. history, archaeology, linguistics, ethography) that any scholar would consider key references, and as few as 100 copies were printed.
The shadow libraries have made a lot available to the whole world. It would suck if private corporations scan and shred remaining copies of these before the shadow libraries can get a scan.
Our little project endeavours to digitise books from the erstwhile Soviet state which were published in many languages
We have, over the last 15 years, acquired/borrowed from libraries, and scanned a couple of thousand books on all topics of interest. All of them are out of print. Some of the physical copies were have, especially in indicate languages might be some of the few surviving ones
Is there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning?
I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense.
> Is there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning?
Is there evidence that they aren’t? All of the common works are already on libgen, if these books are worthless repetition they will contribute little to the training corpus, labs want high quality interesting texts and they have unlimited budget to spend on it. Paying $300 for something rare with millions of tokens of interesting and original text for training is definitely fucking worth it, spending $1,200 to get all the copies and block your competitors from getting it is most definitely worth it!
> I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense.
Do you think it would be OK to rail against for example, a genocide if I wasn’t substantially contributing to some effort to stop it? This sophistic (and uninteresting) bit of rhetoric boils down to: if youre not trying to fix it yourself, dont complain!
(reminder that one of the basic principles of a democratic society is that each person has some concern outside of their personal affairs)
Sure but are we going to do anything sensible about it or is it just a convenient culture war vector?
This is happening because copyright means you can't scan these without destroying them as a format conversion.
No one's felt compelled to try and fix that so we can do this sort of digital archival and preservation, and copyright allows works to be frozen and undistributable because a claim might exist for decades without any actual use (I.e. the number of games which get stuck in legal limbo).
If the only desire is to sling mud at AI companies but not try and improve the legal situation, then it's worse then useless because there's no intent to stop it - in fact stopping it would remove a useful outrage tool.
The idea in the title here is point 1 of the blind leading the blind: it's illegal to scan and store these books without destroying them in many jurisdictions.
>Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.
No, every popular book has been duplicated thousands of times. This is not the same thing as important. They're orthogonal. When an important book is popular, it is safe. When it is not, it is in danger. Only fools assume that important books are recognized often enough to become popular.
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
The other thing they could be doing is using some of that lobbying money to try to reform copyright law to allow them to release the scans that are still covered.
It would likewise earn goodwill from a lot of people.
The United States can't make copyright weaker than what those agreements require without pulling out of the WTO.
The core of copyright law is about who has the right to redistribute a work. If I buy a print of a photograph, scan it and use that as my desktop image... I can do that. I cannot redistribute the scanned image, and if I was to sell the print later I should delete the scanned image.
Note that format shifting is covered under fair use... which is what training is taking place under. However, that doesn't mean that they can release that format shifted content... nor can then re-release the original work if they are retaining the format shifted content.
Under § 106 of the Copyright Law of the United States, the owner of the copyright in a work has the exclusive right to make copies of that work, unless an exception applies. When considering reformatting media, please note that individuals do not have an automatic right to reformat a work from one format to another. In order to legally convert media, your use must fall into one of the following categories:
you own the copyright in the work,
you have permission from the owner of the copyright, or
you have done a fair use analysis and have determined that fair use applies
Here's another question: The exact reasoning for copyright is made clear in the constitution: "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.
Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.
If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?
But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.
Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.
To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?
But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.
I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!
The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.
Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:
a) EU companies making ML models have to self-sabotage against their competition.
b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.
Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of tim...
Ok . How trustworthy is the claim . It maid by a resource that on its own has problems with copyright.
Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge.
It seems the initiative was secretive for a different reason . Buying a single book not allows you to distribute the content of the book , hence 1.5 billions fines.
Usually I personally not on copyright people side , but in this case you clearly see the system functioning. Know knows how the book business will look like in 10 years , but now book publishers doing their job by brining ai companies to court
It is quite disappointing to see them not using the type of machines that don't actually destroy the books, like, afaik, Internet Archive is using.
Using a few gas generators for a transitional period isn't great either, but effect of those is very temporary. I fear permanence in this individual deal with the devil made for speed and cost.
But scanning the books also helps those AI companies because ultimately
they want more data. Yes, they also destroy rare books to sabotage competitors, and thus also damage global society - a reason why these evil companies should be disbanded - but the article seems to not put any thoughts into things here, other than the superficial "they destroy books".
No, it has value because it is the only means that are absolutely accepted to pay taxes and most legal judgements (i.e. obligations to the state that issued that currency.) If you don't have dollars and you need to pay US taxes, you have got to get some or you will go to jail.
How is Anna's Archive getting around the copyright violations of hosting all these books for access to all? I suspect it won't be long before they get sued and are forced to shut down. I spent some time reading the web site, and it doesn't look to be a well thought out project. Even the way it is organized leaved much to be desired. There's much more to library science and the organization of a vast collection of books than meets the eye.
They evade judicial enforcement and try to stay anonymous. They've certainly already been sued (and lost by not even turning up). https://en.wikipedia.org/wiki/Anna%27s_Archive#March_2026_pu... (Contrast to Sci-Hub trying to defend themselves in court, in India, and that leading to no new articles since several years ago.)
Unlike archive.org which is a real business, Anna's Archive is basically piracy.
You could try to sue them but you'd have to find them first.
This IMO makes them more resilient against that sort of thing, and also means they can actually do their job.
the irony of scanning rare books to preserve them while the scanning pipeline is what's destroying the physical copies is going to be a great trivia answer in 50 years
420 comments
[ 0.26 ms ] story [ 70.7 ms ] threadIsn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:
- This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.
- The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.
- The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?
B: Company uses freely available scanned copy of the text:
None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.
I much prefer B.
Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity.
And so can everyone else. That's the ideal outcome.
Only Anti-AI types are against this, they probably don't even care about the books, it's just a proxy for trying to stop the "evil AI companies."
Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.
Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.
So what's the problem here exactly?
Also from the article:
> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
If you've read the articles covering this issue, you'll be aware the concern is over the fate of rare and out of print books, rather than your straw man (ie those available in 'thousands to millions' of copies).
https://www.bbc.com/news/articles/cp3rprx2wl4o
> "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
> "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
> "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
Detestable if they're doing it anyway to prevent competitors getting hold of it.
> "It would be a much more significant problem IF one like that were to be bought for destruction, having survived this long."
I agree with the sentiment of course but it is really a huge IF they are doing that.
IF they wanted to train on, say, Leviathan by Thomas Hobbes, why buy an expensive edition from the 1600s when they would get the same text from a Penguin edition for a fraction of the price? It gets much cheaper secondhand too of course.
I'll go further, why would they want to train on expensive rare and out of print books? Are they, perhaps, competing on an AI benchmark based on extensive medieval knowledge of the cosmos? There's been a lot of pearl-clutching about lost obscure knowledge but y'all really reckon that kind of knowledge is valuable to LLMs?
It’s powerful symbolism. It reminds me of that tone-deaf iPad ad that sparked outrage in 2024. The one where all the cultural artifacts were crushed in an industrial press to make a soulless slab of glass. And that was Apple, who is generally well-liked by the public.
Of course you can. Nobody is saying these companies aren't allowed to do what they're doing.
But what they're doing is disgusting and something I can't forgive. It's an escalation of the attacks against society that these companies have been engaging in from the beginning. This isn't about legality, this is about what's right.
Why on earth would they? Ultimate point for these companies is to make a ton of money, obviously they won't shoot themselves in the foot and give away whatever advantage they have, especially not to a free archive which is about doing good in the world, which probably isn't profitable enough for a company to care about.
How is this a question at all? They’re scanning books because the courts determined that it’s the only way to use that data. They are forbidden from using digital copies found on places like Anna’s Archive. They must acquire and scan the book.
They cannot redistribute the book. The Internet Archive tried that and the courts shut it down. You cannot scan a book and share it without violating copyright law.
Hence the “leaking” part.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of that knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
There was a time of course when you could pull my books down from archive.org as PDFs. Perhaps that time will come again.
I'll look into libgen.
https://gist.github.com/cemerson/043d3b455317d762bb1378aeac3...
> The print original was destroyed. One replaced the other.
So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.
Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).
I’m reminded of the screeds about the dangers of novels.
They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.
A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.
Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.
How tho?
Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take them home, so the seller doesn't have to pack them for the trip back. Just the other day, there was a whole bin of books in front of a shop, offering them for 50 cents a piece. They will be destroyed anyways.
Unless they are buying and destroying really old, rare books or important small-print books, it is not much damage. It is not like they will buy "all copies of all of the books", just one. And its just that the data in physical print most likely hasn't been used for training, so this can help you find more unmined quality data. Nobody is stealing your books, preventing you from buying more or destroying all copies of a single book.
And some of these books would rot out of circulation or be destroyed anyways. Some people throw away 80-100 year old books on the regular, as they might just be unimportant to them or the world in general. And once the last copy is thrown or rots, that book will die forever. This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.
Keep in mind we wouldn't even know this was happening were it not for investigative journalism.
Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.
The internet made this even worse by celebrating people who only have to say the right thing.
The major lesson I myself learned as I got older is that “you should do the right thing” is a rephrasing of “you should do what I want you to do”.
The internet has amplified voices who say things to signal rather than do things to change.
That's orthogonal to knowing what "the right thing" is.
This is as blatantly false as claiming that the Earth is flat, and the fact that there is no such agreement (descriptive moral relativism) has been firmly established in philosophy for well over a century.
Blind futurists and AI cheerleaders scare me on how cavalier and how many crimes against humanity they ignore.
It is pretty standard procedure.
"Hey, guys, when the librarians get pissed about the destruction of books, it’s time to put those listening ears on.
Because we are very comfortable with the idea that books are tools that can be retired. What’s happening right now is not that...."
https://bsky.app/profile/annabookwriter.bsky.social/post/3mt...
I don't think many Liberians like the idea that they have to do it; it is just one of those things that has to be done, sadly.
Weeding (deselection works) is a fundamental part of collections management. Every trained librarian is going to understand this.
The interlibrary loan system has mechanisms in place to make sure the member libraries keep two copies of each work in each region. Collection managers consult these databases during weeding to make sure they don't deaccession the last copy.
If the monograph was never collected by a library and it gets caught up in a destructive scanning project then I guess it was pretty "rare" in a literal sense. "Rare Book" in library land is sort of a term of art and I'm not sure if the books in these destructive scanning projects meet the criteria.
What Amazon is doing is vile."
We're about to enter an era of climate instability that's going to cause wild fluctuations in the ability to grow food across the globe. Historical agriculture data AND data about confounds is crucial for figuring out what strains outside of our current mostly mono-strain agricultural supply chain could be cultivated.
And that's just one use case out of thousands; what if you want to understand and reconstruct technology adoption from that era? What if... you just want to learn what your ancestor was doing at such and such time? What if you want to find clever techniques for robot arms to work with food crops in space?
I don't understand what you're trying to say here.Random article about powder and stone, completely irrelevant.
The old book mentioned was just a farmer's book; you can find all this history in the area's historical records.
"Skill issue" for one is that it's used wrong, but it tells me more that you tend to spend way too much time online and from your insight, very little knowledge about the work of shelters and Library.
If they were simply buying the books and doing nothing with them, the public wouldn’t be able to access their copies anyway.
If they at least didn't destroy the book, someone could purchase it when anthropic eventually goes belly up. Hopefully someone will at least be able to purchase their digital scan and hopefully the scan is of decent quality and clearly indicates the provenance of the text.
Books are considered super high quality training data. Anthropic has no reason to get rid of this data that 1. They’ve spent a ton of money on and 2. Will remain useful indefinitely for training LLMs.
> If they at least didn't destroy the book, someone could purchase it when anthropic eventually goes belly up.
The same applies for a digital scan? The information isn’t any more likely to be lost.
I don't know why you're so quick to assume this will all just "work out" such that the scans are ultimately accessible. Arguably the most likely scenarios are either Anthropic survives and holds them away in perpetuity or Anthropic fails and they are sold off to the highest bidder who does the same.
Yes, there are plenty of books, many were printed, many have lasted a very long time (plenty over 100 years!).
That says more about the success and utility of the technology than it does about whether individual books should be shredded.
But unreadable by humans, right?
You seem perfectly fine living in a world where your flea market is devoid of books.
Also, Anna's text implies that the books are scanned and then mischievously destroyed so nobody has access again to the content. That's not the case: the books are "destroyed" before scanning, by dissasembling them in pages so they can be feed to the scanner. Scanning while keeping the book intact is difficult, as you need to software-unwarp the page before OCR'ing it, and expensive as you either need specialized scanners or humans doing it.
But that's just it, it won't live on forever because, from the perspective of preservation, training is a lossy, noninvertible transformation. The LLM cannot legally produce the book verbatim, it will only spit out a regurgitation of the information, chopped and mingled into a broad information space.
Furthermore, these "magic machines" are not the property of the public. They are owned by a handful of corporations who want to charge you continuously for every token output by the machine. So, not only is the original text locked away forever behind company walls, you now need to pay for access to an approximation of the original contents which you can no longer even verify as being correct because the source is no longer accessible.
If you are cool with this, from a cost perspective you are cool with a deal whereby I trade you access to a definite resource for a one time fee of $N for, instead, a perpetual cost of $M to you every month/day/hour for access to an amalgam in which you cannot even determine what proportion of the resource you are actually getting. You're basically saying you're cool with me selling you some unknown portion of wine for a monthly subscription price instead of selling you a definitive amount of wine for a one time fee. lol.
This sentence reminded me of “The Things” by Peter Watts. The Thing in that short story believes it’s actually the good guy and decides to commit “violent integration” for the sake of humanity.
I’m not sure I buy the links premise that Anthropic is doing this with malice, but if it were, would be an eerie parallel.
If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.
How does this help anything, except create more work to throw in the trash?
It’s called “mandatory deposit”
https://en.wikipedia.org/wiki/Valancourt_Books_v._Garland
> Background Valancourt Books is a print-on-demand independent publishing house specializing in rare and out-of-print books. Valancourt had not registered its books for copyright as the Library of Congress already had original-edition copies of the books Valancourt republishes and any new material in its publications was limited to notes and introductions.
If you want a physical copy of The Sorrows of Satan, you can buy it from them.
Their argument is that the Library of Congress already has a copy of the book ( https://search.catalog.loc.gov/instances/a0f8fcfe-a255-55d2-... ) and having them deposit it again would be unnecessary.
> The Copyright Office has stated that it would modify the language of its deposit demand letters and withdraw its demand for copies if the Copyright Office was notified of the copyright's abandonment.
> Several legislative changes have been proposed to address all elements of the case: changes to Section 407 to tie some legal benefit to the deposit, monetary compensation to copyright holders for depositing books, and regulation for a simple and costless method of copyright abandonment.
That doesn't change that if you were to publish a book today (or for that matter, have published a book in the past 100 years in the US), you are required to deposit a copy of the book with the Library of Congress.
Looked into this a decade ago for publishing eBooks via my personal corp when eReaders and ePub were starting to hit big in the mainstream.
Copyright should really be amended so that once out of print and a grace period it’s free use. I am probably more of an anarchist in this regard. Similar to my belief that anyone should be able to ingest any data you put online, once a book is no longer being print it should be able to be used for commercial or personal use for free. Similar to a generic drugs.
There is far too much garbage that gets published, let the collective hive mind figure out what is valuable.
That's what governments are for.
Sure some governments and opinions would say so but you’re making a statement of zero impact. Fix the underlying copyright laws don’t create more rules.
https://www.copyright.gov/mandatory/
> All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17).
> This law requires two copies of each work published in the United States be deposited with the Copyright Office within three months of publication. Works deposited under this law are for the use of the Library of Congress. Usually, deposited copies must be the “best edition” of the work, which means they must conform to the Library of Congress’s preferred specifications.
> Mandatory deposit applies to any work published in the United States. This requirement does not apply to works first published in a foreign country until they are published in the United States. Copyright registration is optional, but it provides additional legal benefits and fulfills the mandatory deposit requirement with the submission of the required copies.
Just post on r/DataHoarder: "Free 16TB NVMe SSD to anyone who indexes and mirrors the entire out-of-print 20th-century physical archive." The problem would be solved by next Tuesday. With probably 10x redundancy and people willing to do it for free for fun.
https://www.copyright.gov/mandatory/
> All works under copyright protection that are published in the United States are subject to the mandatory deposit provision of the Copyright Act (section 407 of Title 17).
> This law requires two copies of each work published in the United States be deposited with the Copyright Office within three months of publication. Works deposited under this law are for the use of the Library of Congress. Usually, deposited copies must be the “best edition” of the work, which means they must conform to the Library of Congress’s preferred specifications.
----
> Acceptable Formats for Deposit of Electronic Works
> The deposit of electronic works is arranged with the Acquisitions & Deposits division.
> For electronic-only works, submit the best edition in accordance with the formats listed in the “Electronic-Only Works Published in the United States and Available Only Online” section of the Best Edition Statement (PDF, 135 KB).
> For works subject to a grant of special relief, unless otherwise specified, the Library will accept an appropriate “preferred” format listed on the Library of Congress Recommended Formats Statement. Such files must contain no measures (such as digital rights management [DRM] technologies or encryption) that control access to or prevent use of the digital work.
> For more information about electronic deposit, see the above FAQ “When can I make an electronic deposit of a work?”
---
> When can I make an electronic deposit of a work?
> Works may be deposited in a physical format in accordance with the Best Edition Statement, which can be found in Best Edition of Published Copyrighted Works for the Collections of the Library of Congress (Circular 7B) (PDF, 135 KB).
> Works may be deposited electronically in certain circumstances:
> The Copyright Office issues a written demand for an electronic-only book or serial. If your work is published only online and the Office sends you a written demand for mandatory deposit of the work, you must deposit the work electronically.
> The Copyright Office offers you electronic deposit as an alternative to depositing a physical copy of the work. If you receive a letter offering special relief to deposit a work in an electronic format instead of sending physical copies, follow the instructions in the letter or agreement.
---
https://www.loc.gov/preservation/resources/rfs/
https://www.loc.gov/preservation/resources/rfs/text.html
> Neither the deposit requirements of this subsection nor the acquisition provisions of subsection (e) are conditions of copyright protection.
https://www.copyright.gov/title17/92chap4.html#407
Publishing of copyrighted material requires that it be deposited with the Library of Congress.
Copyright does not expire for a very long time. Harry Potter and the Sorcerer's Stone was released ~30 years ago in 1997. It remains protected for the duration of the life of the author (J.K. Rowling) plus 70 years.
Given actuarial tables from the UK[1], this works out to be around ~95 years from now (~2120).
[1] https://www.ons.gov.uk/peoplepopulationandcommunity/birthsde...
The library of congress does distribute some digitized works that are out of copyright. And it does digitize some works for archival and distribution, but having additional works digitized for (eventual) public use could be nice.
It is not at all clear that this is true.
The number of books published every year is growing rapidly. According to Bowker the number of books published every year has increased ~15x in the past two decades [1].
Because of this, I suspect that the median age of these books we are discussing is below 30 years.
[1] https://www.writercosmos.com/blog/how-many-books-published-p...
Let's view it realistically here: AI companies are parasites. Them destroying books to dumb down mankind, absolutely fits into the destruction of the library of Alexandria.
Having said that, I think the day of physical hardcopy of books, is not necessarily over, but will be heavily complemented via digital storage. For instance I only keep books that I may re-read later or read many more times, e. g. thick science books. Many other books I can keep as .pdf file without a problem.
These comparisons are starting to get ridiculous. Why are so many people assuming there is exactly one copy of all of these important books available, and that Anthropic is destroying the lone copy?
Your local library throws out books every year and nobody thought twice about it.
Why are you assuming that each book gets scanned exactly one time and then never again? And why are you assuming that out-of-print books remain easily accessible so long as not every copy has been destroyed?
>Your local library throws out books every year and nobody thought twice about it.
When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.
I know it comes at a shock but truly most books are absolutely worthless.
Many books aren't lent and not bought and most libraries have limited space to store such books, thus they go where old paper goes.
Of course some rarely lent books are important and for the one person asking for it in ten years really valuable, but many still have to go.
I even read on a web site of a librarian that their library had stopped accepting donations because "patrons should know how to throw away their own trash".
How often have you thrown out manuals for products you don't use any more?
After being clickbaited and ragebaited, media consumers feel deeply anxious and angry. But, explaining that they are angry over a boring situation feels silly, not righteous. So, they give summaries, impressions, sometimes extrapolation of the bait they have been consuming. That feels righteous.
This observation applies to a wide variety of topics trending in the various media every day. Distinguishing injustice from ragebait unfortunately requires non-trivial effort from the reader.
That's not their goal or else they wouldn't be burning books. Their goal is making money no matter the cost to the society.
They're destroying one (1) copy of a mass-produced item for each AI company.
Public libraries destroy millions more yearly as a matter of routine.
This is just part of a CCP-aligned moral panic, along with the water use nonsense, and similar with the soviet-aligned moral panic that destroyed the civil nuclear industry 40 years ago.
Notably since all the controversy many data centers are very loud about being closed loop and with significant consideration given to other local impacts as well.
They don't have to, and the vast majority don't.
The lie and moral panic is that all datacenters necessarily waste precious drinking water; it's patently false and used by agitators to push, unwittingly or not, a Chinese Communist Party agenda.
is it actually? it feels more like US propaganda that we have to let our oligarchs run roughshod over us because of what we imagine the big bad CCP might want.
i dont think the CCP cares whether there's data centers in rural america.
Regulation that requires closed loop cooling seems simple enough, same with lots of the other problems people have with data centers:
* sound and infrasound under x DB
* no air quality change
* must pay to build out electrical infrastructure
etc
its not to the CCPs benefit or loss to make sure the data centers are built well if they get built
They certainly care that the US loses the race for AI.
But I understand blaming your energy waifu for its own failure is unacceptable for nuclear bros.
running nuclear plants were shut down while running just fine
What other nonsense do you believe?
Look up Greenpeace Energy, and what cushy corporate job Schroder got after leaving office.
Lying is difficult since you don't have reality to keep your story straight.
I've heard people say this, but I haven't seen any evidence for it. Do you have any evidence?
Exactly the same as shutting down perfectly fine German nuclear plants was aligned with Russian interests.
Not all rare books are valuable. Someone’s self-published junk sitting in the garage is NOT analogous to the library of Alexandria.
Many, most, maybe all of these “rare” books are being scanned instead of just being recycled.
Not a big Reddit fan but there was a great post there from someone in the book industry talking about how non-industry people often give this great moral weight to ever book in a way that is totally disconnected from reality.
All I’ve read, as far as sources go, is a number of rare book sellers saying they’ve had a big uptick in huge orders with no price haggling. Apparently that’s peculiar. And some of them seemed a little concerned.
Now I’m certain they’re not chopping up Davincis notebooks, but I’m not certain there aren’t some that would make people wince.
And I don’t have any reason to think some reddit librarian knows what’s going on, if anything, either way.
As an aside, the entire Google Books corpus is generally estimated at tens of millions of books.
“How dare these companies buy rare and valuable books that nobody else values enough to buy” is a self-canceling argument.
Are you referring to the burning of the Serapeum in AD 391 or the warehouse fires in 48 BC?
Also, helps to clarify what exactly the commenter was referring to and possibly help distinguish the centuries-spanning decline of the Library of Alexandria from the violent fate of the Serapeum.
(But my knowledge of Alexandria extends only to episodes of "COSMOS" and "Connections").
How can anyone say this with a straight face. The knowledge is not destroyed, it is transformed. You can make use of it today in the form of LLMs and the scans still exist. Nothing was lost. It's literally no different from them buying books and stocking them in a private library not open to the public. It's not called the Scanning of Alexandria because if it was, it wouldn't have made a blip in the history, Alexandria's libraries were burned, those books, that knowledge was destroyed. Then only thing being destroyed here is physical copy (again for the people in the back: a copy).
> Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.
Those same copyright restrictions are exactly what would prevent them from sharing the archives. Your beef is with copyright, not the AI companies who are (in this one, rare, instance) following copyright laws/rules.
Hyperbole much?
Does the fact that they're being converted to an immutable digital permanent record for all time mean anything to you? Because as far as I know, the works lost to the Library of Alexandria were wiped out of existence, not simply transformed into a more durable form!
They have done zero to destroy the durability. In fact, it’s probably more durable.
If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.
Are they? Or are they just using it for training and not keeping the digital copy afterwards? And even if they are keeping a digital copy, does that actually matter if they never release it?
You’re making a distinction here, but training is something you repeat for every new point release, so you need to keep the data if you want to use it for training.
Agree. While it certainly isn't the most environmentally friendly to render huge stacks of paper into waste, the real issue is copyright creating scarcity (inability to copy the thing).
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
I ask "Why save physical books?"
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
Doesn't that make them even worse?
That's what I keep saying about the van Goghs I burn to heat my home but everyone is still mad at me!
The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable.
Same can be said about many books.
- Reader: narrative
- Collector: scarcity of the physical artifact
- AI Company: language samples (quantity, variety), facts
"likely" being the keyword here, what about heavily censored books?
Sometimes I don't even know how to respond to comments here. I don't want to be rude, but you just have to give this a moment of thought. Is all the media that you find valuable common? I know that's not the case for me based on my own experience.
The problem isn't with morals or copyright. What we're up in arms about is case law. Past rulings have implied that destruction of books significantly contributes to the process being "transformative". This encourages companies to destroy the books. I think this is really dumb.
Why save physical books? It's because the reason to destroy them isn't good. If you think there's too many bad books out there, that's a different argument. Maybe your fight is against consumerism, I don't know.
That is literally the opposite of how it works.
https://en.wikipedia.org/wiki/Scarcity
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
But where did you hear that they’re buying “all copies”? And to what end?
To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved.
Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.
I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.
Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing ecosystem or the state of a particular political situation at a given time. Essentially all textual primary sources for events which are not currently considered important - but which we have no way of estimating the future importance of. One can continue to create countless counterfactual examples in this vein. My overall point is we cannot know what may be useful or even essential in the future, and knowledge should be preserved under the assumption that it is likely to be.
Historians frequently refer to this paradox - how everyday aspects of life are frequently not explicitly documented, since they're so obvious to the communities or communities of expertise that observe and carry them out. So it's actually incredibly important to preserve what seems like ephemera.
Hell we couldn't have AI training at all if we lacked the corpus of existing written literature - but there was no way any author could have anticipated this future utility more than a couple of decades ago.
1. If someone is acquiring books in bulk for bargain-bin prices and shredding them, they're books whose physical copies have essentially no market value and which, absent this buyer, were overwhelmingly headed for pulping or landfill anyway. Millions of books are destroyed every day.
Could one of these worthless looking books turn out to contain information historians care about in a 100 year? Sure. But that doesn't create an obligation for someone to pay to warehouse every extant copy forever. Physical Preservation has costs: space, cataloguing, handling, transportation etc. Archives and libraries have always had to make choices for this reason.
2. I'm not arguing that preservation has no value. The question is whether destroying a physical copy after digitizing it is a serious loss when talking about mass-produced printed material.
Was this the last surviving copy ? Is the information unavavilable in libraries, archives, other editions, scans, citations, contemporary works etc ? If not, nothing has been lost except one physical instance of a reproducible object.
Your film analogy worked a lot better because old film footage is often unique primary source material. A camera recording of a random brooklyn street in 1993 may literally be the only recording of those people, storefronts and circumstances. The nth printe dcopy of a technical manual is not analogous to that.
Believe it or not, what matters here is the message and access to the message, not the medium.
One could obviously reproduce films too.
Ultimately, you don't know how big N is, and that's what matters. Part of the problem is there is no transparency around this. The population of all books != the population of a specific book or set of books. Assuming N is large when you have basically zero information on specifics is foolish. These books might be just as rare as a film for which only one source exists you don't know.
They're books that are no longer in print. Doesn't mean there aren't a lot of copies around.
>Ultimately, you don't know how big N is, and that's what matters. Part of the problem is there is no transparency around this. The population of all books != the population of a specific book or set of books. Assuming N is large when you have basically zero information on specifics is foolish. These books might be just as rare as a film for which only one source exists you don't know.
All of this is frankly irrelevant. Books that you buy in bulk at barging bin prices are books that have essentially no market value and were going to the pulp or landfill anyway. Nobody has an obligation to spend money to preserve every extant copy of every book forever. Millions of books are destroyed everyday. Even libraries and archives make these choices. I was simply pointing out the film analogy worked better, but it doesn't change anything.
You realize "market value" isn't the only kind of value too, right? Even books that are completely obsolete can have inherent historical value insofar as they enrich our understanding of the past and how human beings used to live. Contemporary market value isn't everything. There are plenty of works of scholarly interest that wouldn't command nearly requisite value to a NYT bestseller when you adjust for scale. That doesn't mean these books aren't worth keeping or aren't intellectually important.
Also, these books presumably even have some perceived value, otherwise these companies wouldn't think they were relevant for training, and wouldn't be spending money to purchase and scan them in the first place. The whole premise is that these books are valuable enough such that training on them will make the models valuable and make paying customers line up for tokens. This is an inherent imputation of value to these works on the part of these companies. You can't say, simultaneously "these books are worthless" yet somehow, at the same time, "they are worth something for producing artificial intelligence". Your assumptions don't make any sense.
> They're books that are no longer in print. Doesn't mean there aren't a lot of copies around.
Sure, but there also could be limited copies around. We don't know. We won't know unless these companies are more transparent about exactly what they are buying and scanning. It's that simple. If it's foolish to assume there might be limited copies, it's equally foolish to assume they have plentiful copies. We simply don't know.
> Nobody has an obligation to spend money to preserve every extant copy of every book forever.
I don't think anyone is claiming that. People are bothered because they (in my mind, rightly) recognize that old books are a historical record of human accomplishment and have some amount of cultural value. Letting a handful of companies further plunder the collective output of humanity and now, potentially keep it locked on their servers in perpetuity is something that should make you upset if you have any inkling of interest in the history of humanity or human accomplishment and if you have any sliver of curiosity about these things. If all you care about is the present and getting better LLMs, sure, I guess I understand why you don't care, but then I also think you are very unwise and myopic.
> Millions of books are destroyed everyday.
So what? Millions of crimes are committed every day, that doesn't mean we shouldn't be bothered by the crimes we hear about. Billions of pollutants are spewed into the environment every day. That doesn't mean we shouldn't try to combat pollution. That's a pathetic attempt at justification. Also, there are clear distinctions between processes like libraries removing works they can no longer store and companies scooping up books for training and personal gain.
https://downtownbrown.substack.com/p/five-fallacies-ai-and-d...
It's not that individual companies buy all copies of a given book, but that there's more than one book scanning company, and they aren't sharing the scans with each other. The result: books that were rare but nevertheless easy to find for purchase (thanks to the internet) are now vanishing off of the market, becoming de facto no longer accessible to the public.
If anyone has an example, I’d love to hear it.
The distributed nature of the problem and the utter lack of transparency are huge factors here too.
Regardless I feel like they gotta figure this out for optics reasons. “So and so books are lost forever to Anthropic’s servers” is much more outrageous than “Anthropic is destroying a bunch of books that have other copies” imo.
Thank you anthropic! Thank you OpenAI! Thank you Google!
somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”
the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.
> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.
archivists keep everything, because we don't know right now what will be important 100 years from now.
By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther.
Plus having the info part of a LLM makes it immediately available to literally billions.
I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.
Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong.
If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions".
And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.
For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book.
The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly.
But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great.
That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.
No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference.
> So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of.
That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy.
I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years.
I understand what you mean but... "permanently lost" sounds dramatic. When I trow away old pictures, old drawings or pieces made by my son at school, they are lost as well. Not that I do that often, but it begs the question, should every 'ip' made by humans be preserved?
A small fraction of them is saved in the model. Far more is saved in the digitized copy as long as they keep it which they have plenty of incentives to do so (future training of newer models).
That's more than what happens if that book was burned or sent to a landfill, but less than if the book is giving a loving home.
>They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine.
My understanding is that this simply isn't legally allowed for these books. The original must be destroyed for the digital copy to not be copyright infringement.
>That’s a false dichotomy.
I pointed out there is a spread of possible outcomes and that different people are considering different outcomes and the comparison of if this is good or bad depends upon which outcome one considers. I even mention that both outcomes are sometimes right. That's about as far from a false dichotomy as I can see it.
> Plus having the info part of a LLM makes it immediately available to literally billions.
Isn't this a contradiction? I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions".
I don't if that is true: a lot of old books might still have copy or other rights associated to them, likely owned by author and/or publisher, directly or inherited, but often those who have the rights do not have digital or physical copies at hand anymore (some old books are, well, really old). Does Anthropic make sure to track down, contact and then share the digital copy they make with those who have rights on the work? If not, they are not making it in any way easier to re-print the books, while making their supply more scarce (they destroy existing embodiments).
Citizen's lobbying against megacorp is a mirage.
>But it's _closer_ to being widely available, not farther.
By what metric ? The copy is now guarded by a company instead of being on the second hand market.
Destroying the last copy of a book is OK because you shop for second hand books, and because some private company hold the last digital copy and have no incentive to make sure it survives. Damn..
History tells us that very few "permanent" situations are truly permanent.
Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.
If you destroy the only copy of a physical artifact, the situation is as permanent as it can get.
So physical books are rather non-permanent themselves.
Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway.
Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why?
[0] https://www.bodleian.ox.ac.uk/services/research-partnerships...
Do you think there's no difference between the books curated at Oxford University and the crap that Amazon is buying?
I have volunteered with a library, and probably threw hundreds of books over a couple week engagement from a university library into a shredder at the direction of a professional, academic librarian. Libraries are constantly culling books, the EXACT category books we're talking about here (old, never-read). This is happening at a much larger scale, so I would recommend railing against university librarians in addition to the AI juggernauts.
Have you seen evidence that they’re buying only readily available books that are plentiful on the market?
A few years ago, after running into this issue several times, I looked into the economics of it to see if there was a opportunity to republish digitally. I found the sales data for some, the ones that had value were selling on average for around $50-150 dollars. Because the typefaces weren't modern, OCR wasn't scalable. Because only 40-50 were transacted each year, it wasn't worth the time.
It's interesting that libraries are purging them. Several times, I've found that the only available copies were at some random university rare book collection in middle of no where, and basically impossible to access unless one fly in - which isn't worth it. They are often donated by a benefactor and stuck there - which is why they're never read. It's not that the content isn't valuable.
What kind of books are they? In my case mostly historical documents by some relatively unimportant person who was highly important for a very, very niche subject.
They still contain valuable, irreplacable information. And, once the physical copies are destroyed, they'll probably be gone forever. An analogy: imagine you discover a really cool video game from 20 years ago. You love it and want to find the developer's previous work. But, you find out that the company was purchase by another company which was purchase by another, etc, etc. Sure, maybe the game still exists in some digital vault. But, more than likely it's gone because old things only seem have value these days if some influencer advertises it.
The books we were disposing of were so totally inane that I couldn’t tell you what they were, but they weren’t taken out in like 30-40 years. Northeast US state R1 research university. Maybe we got rid of stuff that would come in handy!
Fetishizing books isn't going to help.
In fact, most of those "rare" books don't sell because nobody wants them. The AI companies are making them even more rare, so the booksellers should be thankful.
Books in the former USSR display their print runs on the last page. "Important book" is a vague and arbitrary term, but rhere are works in whole fields (e.g. history, archaeology, linguistics, ethography) that any scholar would consider key references, and as few as 100 copies were printed.
The shadow libraries have made a lot available to the whole world. It would suck if private corporations scan and shred remaining copies of these before the shadow libraries can get a scan.
We have, over the last 15 years, acquired/borrowed from libraries, and scanned a couple of thousand books on all topics of interest. All of them are out of print. Some of the physical copies were have, especially in indicate languages might be some of the few surviving ones
https://mirtitles.org/
https://archive.org/details/@mir-titles
It is common for academic books to have publication runs in the low three digits.
You may argue these books are not important. But how do we know if we fail to preserve it?
I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense.
Is there evidence that they aren’t? All of the common works are already on libgen, if these books are worthless repetition they will contribute little to the training corpus, labs want high quality interesting texts and they have unlimited budget to spend on it. Paying $300 for something rare with millions of tokens of interesting and original text for training is definitely fucking worth it, spending $1,200 to get all the copies and block your competitors from getting it is most definitely worth it!
> I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense.
Do you think it would be OK to rail against for example, a genocide if I wasn’t substantially contributing to some effort to stop it? This sophistic (and uninteresting) bit of rhetoric boils down to: if youre not trying to fix it yourself, dont complain!
(reminder that one of the basic principles of a democratic society is that each person has some concern outside of their personal affairs)
This is happening because copyright means you can't scan these without destroying them as a format conversion.
No one's felt compelled to try and fix that so we can do this sort of digital archival and preservation, and copyright allows works to be frozen and undistributable because a claim might exist for decades without any actual use (I.e. the number of games which get stuck in legal limbo).
If the only desire is to sling mud at AI companies but not try and improve the legal situation, then it's worse then useless because there's no intent to stop it - in fact stopping it would remove a useful outrage tool.
The idea in the title here is point 1 of the blind leading the blind: it's illegal to scan and store these books without destroying them in many jurisdictions.
have you ever held and read an old book?
No, every popular book has been duplicated thousands of times. This is not the same thing as important. They're orthogonal. When an important book is popular, it is safe. When it is not, it is in danger. Only fools assume that important books are recognized often enough to become popular.
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
It would likewise earn goodwill from a lot of people.
Berne Convention https://en.wikipedia.org/wiki/Berne_Convention (182 parties)
TRIPS Agreement https://en.wikipedia.org/wiki/TRIPS_Agreement (164 parties - part of WTO)
The United States can't make copyright weaker than what those agreements require without pulling out of the WTO.
The core of copyright law is about who has the right to redistribute a work. If I buy a print of a photograph, scan it and use that as my desktop image... I can do that. I cannot redistribute the scanned image, and if I was to sell the print later I should delete the scanned image.
Note that format shifting is covered under fair use... which is what training is taking place under. However, that doesn't mean that they can release that format shifted content... nor can then re-release the original work if they are retaining the format shifted content.
https://library.georgetown.edu/copyright/fair-use-reformatti...
If their PR teams did have a response to these actions, it would be something to the effect of:
"would you rather china destroy all the books and gatekeep the knowledge?"
What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.
Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.
If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?
But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.
Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.
To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?
But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.
I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!
The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.
Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:
a) EU companies making ML models have to self-sabotage against their competition.
b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.
Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of tim...
Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge.
It seems the initiative was secretive for a different reason . Buying a single book not allows you to distribute the content of the book , hence 1.5 billions fines.
Usually I personally not on copyright people side , but in this case you clearly see the system functioning. Know knows how the book business will look like in 10 years , but now book publishers doing their job by brining ai companies to court
Using a few gas generators for a transitional period isn't great either, but effect of those is very temporary. I fear permanence in this individual deal with the devil made for speed and cost.
It's not bitcoin.