62 comments

[ 0.20 ms ] story [ 20.1 ms ] thread
you do realize "rare books" in this context most likely means random technical manuals nobody cares about and not collectors items, right?
Unless we know this to be the case it's reasonable to assume it might be more.
Oh, that's really interesting, I'd love to see the list of books that they've digitized too. Where did you find it?
follow the discussion around this on X, i think just a few minutes of research on this topic / reading past the headline you'll find out that this is pretty much a nothingburger
I'm less concerned with how these rare books didn't rot in large lots of unused books and more concerned with whether or not this helps preserve the content longer.
The content is preserved obviously, and I hope that sometime in the future it will be made available in its original form.

What worries me about the trends is inevitable sanitization of content or straight out falsification.

... because now it can be done at scale.
>What worries me about the trends is inevitable sanitization of content or straight out falsification.

That sounds really speculative, and not inevitable at all.

I think you're just trying to invent things to be worried about because you don't like AI and don't trust AI companies.

Well, 'history is written by the victors' is not just a toss.
I seriously doubt any of the companies doing this will preserve the scans for very long. It costs money to store stuff, and none of these companies are have any concern for anyone who isn't them, so they're not going to spend the money or lift a finger unless they work out a way to make it profitable.
archive.org estimates it costs $2/GB to store data in perpetuity (well, at least until decades of storage scaling trends stop of course). That's probably less than it costs to acquire and scan the data in, I'd be surprised if they just tossed it out at the end when it's so cheap to store. The same thing happened at the healthcare data lakes I worked on where the trend switched from the usual "how long are we legally required to store this information" to "how long can we legally hold this information".
I remember when hackers believed in the doctrine of first sale. You can do whatever with the stuff you own.
They are rare because nobody cares about them otherwise. Why is everyone acting as if they are trashing Gutenberg Bibles or first edition LOTR copies?
I really have trouble getting worked up about this. "Rare books" is thrown around regularly but my gut feeling is that's not the case. These are used (often? always?) books and while I'm sure there is waste, in general they just want 1 of every book.

While I wish there was a repository of every book that was already digitized (it pains me this is the best solution), there isn't one and so I think this is not a real problem.

It'd be a different story if they had furnaces that ran only on rare books that they had to continually feed books to but that's not what's happening here. And that 1 destroyed copy will "live on" in a way that it otherwise might not.

It'll live on if they publish those scans or contribute them to a national archives or something. Proprietary data has a habit of being lost over time though.
The problem is that we're ignorant of the true value of objects and we don't know what will be valuable in the future: https://en.wikipedia.org/wiki/Palimpsest

It seems a little short-sighted to destroy an artifact to get the text.

Perhaps the genetic material that remains in books from the people who handled them will have value in the future, but we won't know what we lost because some people foolishly destroyed it in a bizarre quest to make AGI that the creators argue could potentially destroy humanity.

The more and more I read about these kinds of people the more I'm starting to realize that they're in the "here for a good time not a long time" group of people and those are the last people you want making long-term decisions.

Millions of books go in the trash every day. Should all private property come with such an asterisk that it ought to be preserved for whatever unlikely contingency you can imagine? I don't know, AI actually seems like a much more worthwhile pursuit.
On its own that seems like a pretty weak justification, at that point we would need to preserve every copy of every book forever because you never know what what useful physical material it has on it even beyond its text. That extends to non-book items as well, and while I understand the idea it's just not realistic in any sense.

I think the more important part is to define what "rare" actually means and save/digitize set copies of those books that are actually at risk of being lost completely (rather than letting them rot away somewhere or get bought up to be privately destroyed). I suspect many of them are just not at all interesting enough to justify the expensive though.

My comments on this subject aren't some sort of justification for any particular course of action, they're more a disapproval of the course of action being made by people people who have more dollars than sense.

As a general rule preserving historical artifacts for continued future analysis and appreciation of by people who have yet to be born is a noble cause that's considered worthy in and of itself without the need to justify it. We're talking about the richest group of people who have ever existed with the technological means to do but decline to do so and instead act like the Taliban blowing up statues of Buddha.

I think this is intuitive to most people. If a prerequisite for the birth of AGI was that a humanoid robot had to sit down in the Louvre and eat every single painting there with a knife and fork while occasionally stopping to wipe away historical detritus with the Mona Lisa that it's wearing as a bib people would by and large express a visceral and justifiable outrage.

The people who are doing this know how this looks so they're trying to do it all behind closed doors, through cut-outs and intermediaries.

They are literally not allowed to publish those scans.

Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright from strangling all the freedom in this narrow area.

If they are in the public domain, they could publish them. How many books are in the public domain but have never been digitized and shared in a public archive?

And if the books are not in the public domain, then they should not be allowed to train their AI models with the material without some kind of license or agreement with the owner of the copyright.

There's no public domain from an international perspective. After the Authors Guild settlement Google Books took down alot of Irish and English case law from the 18th and 19th centuries because, AFAICT (assuming it was even given serious consideration), there were some potential esoteric copyright or publisher rights they didn't want to deal with. Even some old American legal literature disappeared, perhaps because some of the material remains in continuously updated legal treatises and Google didn't want to deal with the headaches.

Many countries claim exclusive publishing rights over even ancient works. If you publish anything online, no matter how old or esoteric, you're taking a risk of violating some kind of copy, publisher, or artistic right. The risk might be miniscule, but it's there. If you have an international footprint, the risk grows considerably; you can't just ignore claims from a jurisdiction you have a presence in or may have in the future.

> If they are in the public domain, they could publish them.

That'd be an easy fix with new law -- if you are an AI company with book data, you have a burden to make openly available (or require your suppliers to) all public domain book scans.

> They are literally not allowed to publish those scans.

Emphasis mine. Just provide a digital copy, no questions asked. They will distributed to various archives globally. I'll pay for the drives and shipping. I understand and can appreciate the potential liability, and am willing to launder it to preserve the subject collection(s) and dataset(s).

Rare books tend to be out of copyright, true.
We run a little library in front of our house. It's amazing how many people dump boxes of old books off in front of it hoping that they will find their way back into someone's collection.

The sad truth is the vast, vast majority of printed literature is neither interesting nor useful. People are not dumping off stacks of Umberto Eco. We frankly have to toss a lot of awful cookbooks, self-help books, trashy mass-market "novels", and sketchy religious works. As it is, even the stuff that makes it to the library is not very impressive.

Imagine if we had bad cookbooks, cheap popular novels and stories, tracts on diet and self help from old Rome, or old eras in China, or the equivalent from ages before that.

It's all interesting for something, even if it's just a meta analysis of culture during a certain period or what kind of trashy romance novels were popular in 198X. At least in my view.

> Imagine if we had bad cookbooks, cheap popular novels and stories, tracts on diet and self help from old Rome, or old eras in China, or the equivalent from ages before that.

It would be great, but it's a bit off topic here. The point the parent is making is that most of these works are going to be trashed regardless of Amazon's behavior. If Amazon (or any other company) is digitizing it, they're at least preserving it in some form. The alternative may well be that many of these books never get preserved.

Now would I prefer the government or an entity like The Internet Archive do this? Sure.

They were all lost when the library at Alexandria burned, so clearly we should ban libraries. Too much risk of losing unimaginably priceless works all gathered in one place like that
> they just want 1 of every book.

That's the point I come to.

Books are not original manuscripts. Even in low volume cases, they are usually printed hundreds of times. (And usually low volume works aren't all that great...hence the low demand.)

That is a fraction of a percent for books that at some level weren't all that wanted.

Do we know if they actually buy one copy of each book here? Or are they indiscriminately buying books in bulk and just processing all of them?

The latter seems inefficient, so my first assumption would be that they would avoid that. But while I suspect they'd check if they know the book before scanning, I could imagine them not caring that much before buying them and just focus on volume.

Agree. Couldn't care less. There are some neat aspects to "rare books" but overall quite insignificant.
I remember when google was scanning a bunch of rare books, I mean they might still be doing that? Either way, that was cool.

I have a few "rare books" and have read many, you'd be suprised at what is publicly available on google books since like ~2010ish.

The scariest thing for me about companies not caring even minimally about conservation is that when AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here), I want to hope that AI will care more about conservation of human people.

So far from how I see how powerful organizations work, I'm not as certain as I would like to be.

> AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here)

People say this all the time, but so far nothing has convinced me it's true.

LLM development has more or less plateaued, and the current boundaries are very real - energy, resources, capital.

At this point we're talking about marginal improvements against the same asymptotes of all technological innovations.

A couple thoughts:

  * They're only buying a single copy of that book as they only need one to scan

  * If the book was public domain (or should be), then there should be an effort to "democratize" that data into a public commons of intellectual property?
Would that an enhancement to the Library of Congress or such?
A couple answers…

Why do you think for even a moment that they would buy only one copy, rather than every available copy of something rare and difficult to get? Anyone doing this stands to gain from destroying the only copies of something that has scarcity, to stop competitors from ever getting access to it. That's a serious concern.

Why would they do that? What is this 'should' to a company that would do this in the first place? Are you not suggesting the opposite of what they aim to produce?

That seems really speculative. I think you're assuming the maximum possible evil intent for no reason, mostly because you hate AI and anyone involved in it.
They are only destroying the books because they are required to by copyright law. They obviously wouldn't do so if they were allowed to merely copy the book and preserve the original.
No, they aren't, for training AI, at least not based on anything but pure speculation.

The recent trial court decision that keeps being pointed to to support that:

(1) Found that for training AI, digitizing and copying works was fair use, period, with no requirement to destroy.

(2) For creating a centralized digital library for general use, digitizing works while destroying the hardcopy was fair use even with no intent to use them for training AI.

Am I too old now, expecting someone to make a Rainbows End reference? Vernor Vinge predicted this 20 years ago.

(Also the person who coined Singularity, though Ray Kurzweil really wanted everyone to think it was his idea.)

Yep, unpleasantly close prediction. Probably shouldn't give them any ideas.
That is certainly an interesting case for thinking things but not typing or saying them out loud.

Problem is, they all read the same books we do.

This feels like a manufactured controversy. What difference does it make to me what someone does with a book after they buy it? It's effectively unavailable to me regardless of what they do. If people are really concerned about these "rare" books, they should lobby the copyright owners to release them online or print more copies.
What difference does it make to me if someone shoots the last bison? I wasn't getting to eat it either way.

If people are really concerned about bison, they should lobby gamekeepers to release photos of them.

https://en.wikipedia.org/wiki/American_bison#/media/File:Bis...

A digital copy of a book is identical in value to a printed copy.
This is entirely untrue.

Simple example:

Actual book from 1732 (rare, original, older than your country): https://blackwells.co.uk/bookshop/product/The-Compleat-City-... -- yours for £1,258.00

A scan of the same edition of that book: https://archive.org/details/bim_eighteenth-century_the-compl... -- free. Nobody's paying any money for the digital copy.

I don't think you understand the book itself is a collectible object with rarity and value, regardless of the information it contains.

Plus, the physical copy of the book has additional value in terms of forensic verifiability. You can prove the originality of the text and the absence of alterations, and can easily track changes between editions.

Digital copies can be altered at whim, as the only means to provide any sort of tracking or verifiability is through another external software system that itself has to be trusted.

Replace bison with some random animal no one's ever heard about, and photos with literal clones of the animal, and you've got a much better metaphor. Not to forget that in this world, brand new animals are created every day.
> What difference does it make to me what someone does with a book after they buy it?

I think this is a bit of a myopic take. It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.

I'm with you on the "rare" part. If people are thinking about 70+ year old documents or ancient manuscripts, I doubt that's what AI is being trained on and is being destroyed, but it's reasonable that people find _that_ idea distasteful.

You can say it's manufactured but if these companies ignore this criticism, it's just another way AI companies are committed to losing the public.

> It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.

Yes, and when you buy them with that designation, you know what you're getting into. The problem occurs when you already own it, and some group is trying to get it labeled as historic, which will add to your burden and limit what you can do with it.

When I was in a small town, this was actually weaponized. A hotel owner was trying to get another hotel categorized as "historic" and had rallied a lot of people behind his cause. He had a case - the hotel did have some claim to being the "first" in some category or other. But really, he was doing it because it was a competitor. The "historic" hotel owner had to spend a lot of money to fight the cause, because being labeled historic would prevent him from performing various upgrades, making the hotel less attractive to customers (he was already not getting many customers).

(comment deleted)
More and more, I am glad that I stopped giving Amazon my (formerly) enormous amount of business.
>It'll live on if they publish those scans or contribute them to a national archives or something

I'm not really aware of many benevolent acts Amazon has taken in the past decade. Are you?

So, if a corporation can suck up entire books to teach their machines how to think using the information from those books, can we (all humans) join a single corporation that provides all books to its employees? Just need one copy of each and we'll make that copy digitally available for our employees so they can learn from and utilize the knowledge from the books. It's not copyright infringement, they're employees.
Almost like… a library?