41 comments

[ 0.26 ms ] story [ 13.4 ms ] thread
This is what the copyright laws dictate no ?
No. A copy is still a copy even if you destroy the original.
Until about a year ago this would have been a reasonable and respectable argument, but this year, at least in California, you are arguing against legal precedent: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
That's a bit of a wild interpretation of copyright law.

I mean, Anthropic isn't going to fight it because it lets them do the thing they want to do, so I can see how this never gets beyond the court that allows them to do the thing they want to do.

But would this argument would have flown in the past?

It wasn't even attempted in Sony v Universal. Or any copyright suit up until this point. That doesn't smell funny to you?

Sure, it's a bit wild, and no, the same outcome may not have been reached if the case came up earlier. And if it makes it to the Supreme Court, who knows what they will turn it into.

The flip side is that Alsup (the judge who wrote the opinion) is probably the smartest district court judge we have when it comes to technology, and one of the people I'd trust most to come up good decision.

He's a treasure, and my instinct is that he got it right: https://en.wikipedia.org/wiki/William_Alsup

i would assume that would come into play if they were uploading scans of the books? there must be some gray area where training like this doesnt apply to that. or ya know, just do it and face the consequences later because you already have the data and know that ai obsessed government will just shrug their shoulders.
Not to my understanding. To begin with, it's far from given that "rare" books are all covered by copyright. But if they are, it's at best murky: whether you destroy the original doesn't really have anything to do with what you're doing with scanned contents. The scanned contents themselves may be inherently a copyright issue, regardless of destroying the original. The actual trained model has separate arguments more in its favor, so if no scanned contents exist - IE the data is read once for training and not stored or saved, they have a better argument. But in that case the destruction is totally disconnected from copyright, as they'd be totally okay to rescan the material.
Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.

Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.

If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).

They could just not do the evil thing.

(This is why I will never be a billionaire)

Those laws are lobbied for by large corporations, these are not just laws that exist outside of that context. They can also be changed, or Amazon could just incur the fines.

Large corporations will move fast and break things when it’s convenient; they don’t care much about the law - just about profit.

TechCrunch is trying too hard to make a connection in the title there. Amazon was trying to make money then as it is now. Their selling of books at the beginning was no more principled than the destruction of books now. This is the legal loophole that allows them to do what they need to do and the benefit of doing it outweighs the modicum of outrage this title will generate.
The worst part of Amazon destroying priceless old books in the quest to build the torment nexus is the hypocrisy
all due respect but if they buy them they can do whatever they want with them. not like they are people and not like you have a right to them either. not even like you would have been able to acquire them yourself.
You're not really responding to the post, which is about the hypocrisy.

"They're allowed to be a hypocrite" doesn't mean they aren't a hypocrite.

Okay, let's say some books deserve the equivalent of UNESCO status. How would you, personally, go about this?
They obviously aren’t priceless if Amazon is buying them for pennies.
Is it a “loophole” that you can buy something and burn it? That’s how property ownership works. If the books were really so rare and valuable then the original owner should have given them to a museum rather than sell them as scrap. Yet no one who is complaining right now cared about them before Amazon got involved.
No, that's not the loophole. The loophole is that a US federal judge recently ruled that scanning a book for the purpose of training an AI is "fair use" and hence legal _if and only if_ the book is destroyed in the process. If you retain the original, you would be creating an unauthorized copy, but if the original is destroyed, then it is permissable as format shift: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
Sometimes original owner does not know that it is a first edition single copy. For them it is an old book. There have been many cases where owners did not know true value of their antiques.
Well in this case the original owner doesn’t know, the purchaser doesn’t know, the reporter doesn’t know, you or I don’t know. So what actually makes these books rare?
> Is it a “loophole” that you can buy something and burn it? That’s how property ownership works.

It sounds shocking if you're missing the context and didn't go past the blogspam which is this TechCrunch article.

They aren't buying the books only to destroy them but to build an AI training dataset, which means they're making a digital copy. This becomes a copyright issue. The destruction is the legal loophole that allows them to keep the digital copy as the only copy in circulation and be considered fair use as decided by a judge (see below).

The concept of ownership means you can do anything you want with the object, the book in this case. Not with the content, like make or distribute copies. The scanning machines are cutting the spine of the book, feeding the pages to scanners, and then destroying the physical copy.

Anthropic's version of this was Project Panama [1]. Quoting from the page:

> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use

[1] https://en.wikipedia.org/wiki/Project_Panama

Its somewhat ironic.
There's nothing in this article supporting the claim that rare books are being destroyed, it just has a link to a paywalled article from 404 Media.
"rare" is used in these headlines/articles to incite and generate clicks
Tech crunch used to be tech news, now it’s ai slop with a biased agenda.
no one uses ai for writing at tc. you're gonna end up chicken-littling yourself
It used to be breathless startup hype and gossip, which it might still be in some capacity.
The Embassy of the Free Mind (https://www.embassyofthefreemind.com) is a rare book library in Amsterdam that is scanning books the old fashioned way… leading to https://SourceLibrary.org — a collection of over 5,000 books from the renaissance that have never been translated before. Consider donating, if this is a topic you care about!
Rare as in your grandfather's John Deere manual from 1982, not rare as in a test print run of The Great Gatsby. The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
> The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.

Citations? As a bibliophile who loves scouring used book stores for out-of-print titles this is a topic I'm very interested in.

> Also what exactly are these 'normal means'?

Normal means is throwing it in the recycle bin. Especially for stuff like a 1982 John Deere manual. I've never donated an old appliance's manual to the library. Have you?

Amongst my friends, I'm one of the rare folks who donates books to the library. Most people just trash them. And I know the library only wants them to try to sell them in their book sales (or online) so they can get money. Almost nothing one donates to a library actually ends up on the library shelves.

Carnegie was building libraries and concert halls, Silicon Valley CEOs buy girlfriends with big tits, cheat in computer games and destroy books.
In the end its not who can train the better model its who owns the better stack of training data. Acquiring old texts then destroying them makes their information proprietary.
So... do the materials that it is printed on make it rare? Or is it the combination of ideas, paragraphs, and words that comprise the text make it rare and therefore valuable? If they're releasing digital versions of this text, i feel that is better off anyway.
Well, I hate to be the one to break it to you but they're not releasing digital versions, they're just destroying them.
The headline implies at some level that Amazon loves or cares about ..... stuff.
Well they have to destroy them after scanning, else it is a copyright violation. Maybe not in the US, but in a lot of jurisdictions you are allowed to make copies/scans of Books, CDs etc. but if you sell the media, you have to delete the copy.

Given that anthropic was just fined a $1bn fee for torrenting/piracy, we cant really blame the companies for shredding the books after a scan. We should blame copyright laws.