49 comments

[ 3.0 ms ] story [ 64.0 ms ] thread
One of my hopes is that when the AI bubble bursts, some brave person will sneak out a copy of the last frontier model.
Piracy / copyright predictions?

The current situation feels untenable with renting. So many regular people I know have learned about VPN, NAS, etc.

[dead]
Some more interesting bounties they offer: https://software.annas-archive.gl/AnnaArchivist/annas-archiv...

> Purchase all Library of Congress MARC datasets — $3,000 bounty

> English Wikipedia pages about relevant institutions — up to $100 per new page

> Internet Archive Digital Lending — $5000 per 1 million pdf files

> Text version of our full library — $20,000

...

(comment deleted)
Who is behind Annas archive, there is a lot of english speakers involved in the team and forums! Anyway as long as buying isn´t owning no issues here.
I wonder how long it will be before they offer bounties for internet scrapes.

Cloudflare captchas have made the internet unusable for me, and I'm sure it will only get worse over time. I'd much rather just browse (or even torrent) a copy of archive.is or similar. The latter would be much better for privacy, and hey, I run ad blockers anyway.

Curious as to how you would approach this. I have no experience in this area, anyone on this forum willing to share their expertise?
The US should just find a way to quietly share literature access with the Russians, rather than letting piracy be promoted and facilitated for US consumers as freedom-fighter "archiving".

Between all the piracy, and all the AI training and the purchase/visitor-circumventing AI services, the practice of writing and publishing genuinely good work is being wiped out.

We're killing the goose that lays the eggs, for selfish gain.

Do you have stats on that?

I’m not sure piracy or AI training are really affecting book publishing dramatically. But if you have data, I’d be curious to see it. AI scraper bots are a total pain for online publishers and FOSS sites, but AFAIK they’re not really harming book publishing directly.

The consolidation of publishers and Amazon’s own practices are probably worse for authors than “piracy”.

Russians will just share it back (I’m saying that as a Russian). And if not Russians, then somebody else will.

What you can do is make sure people can pay you easily, and not put (a lot of) hurdles in your readers way. And when people can’t afford to pay... maybe let them enjoy your work still, and you’ll get a couple more loyal fans who would pay you when they’re able to.

At least this was my world view before AI has arrived and ruined^W disrupted everything. Now I’m not so sure.

AI publishing is just email spam, but for books. When the cost of creating worthless text is low, people do it.
I live in a country where the selection of available books, especially in English, is very limited. Buying online from foreign markets comes with a long list of administrative hurdles and limits.

If it were not for Anna's Archive and Z-Library, I would've never been able to read the books that shaped who I am today, or keep my passion for learning alive.

Thanks, AA and ZLib! (Also, thank you to the authors whose books and knowledge I consumed without being able to pay them back.)

I can feel for both camps.

Some of my published work is pirated heavily. That's not my main income source, so I just shrug and let it go. If anything, I'm probably happy that people are reading my work. Especially if it's people that can't afford it, I'm glad they enjoy my work -- those books got pretty high reader ratings, and it seems to me many readers are actually reading the pirated version.

But I do have friends that depend on this income source, and fighting piracy has become a part of their day job. It's not a fun thing to do, they'd rather spend time working on their next story, but they still have to do this everyday. I feel for them.

That's almost every country these days. At least in the EU. Amazon doesn't carry nearly the selection you'd want, and certainly not in the non-fiction department.
(comment deleted)
https://SourceLibrary.org has about 16,000 rare books translated — most for the first time. 50,000 books archived (will be translated when we have $$ for it). More tokens than English Wikipedia and about .75 petabytes.

Not sure if we will qualify for a bounty, but happy to share! Btw, we are looking for funding from small or large donors who want to help us translate the Renaissance…

The link sort of reads like people who have very easy access to the requested material. Almost like they're Google employees.
Does Anna's Archive use a completely different "source repository" from LibGen?
The only legal hurdle keeping Anna’s Archive away from its noble goal (piracy laws) has been shown to mean zilch in the age of AI.
There was a time where you would get a random page preview, some artists found a way to extract full books that way (F.A.T lab?).
How is Anna's Archive funded? I see they have memberships, but it's hard to believe that can fund all these bounties - some going into six figures. Ask any FOSS project about funding by that method.

It seems like there are some deep pockets funding them.

Just do it and be legends, Larry. ;)
I think this would cross the line from civil copyright claims into criminal activity

https://chatgpt.com/share/6a4970e8-7fe8-83e9-8f81-3aefd76b6b...

On another note, if Google's cybersecurity were always one rogue employee away from a massive leak, then it wouldn't be Google. What was the last Google leak you remember, defense in depth people.

AA is an openly criminal organisation. Their attitude to prosecution is "you'll never catch us lol"
Anna’s came clutch for me yesterday. I spent a few days trying to find a zip file of a CD that came with an old book from early 2000s on programming. One of those Thomson Publishing slap jobs that I actually enjoyed. I checked used copies all of them said does not come with CD. I tried googling around, nothing. LLMs couldn’t find it. ChatGPT kept saying it is on the archive (no it isn’t you useless piece of shit). Anyway, on a whim I went to AA, lo and behold, zip files for both first and second edition. Godsend.
Consider putting them on archive.org then, they have a section for that and it helps spread sources

    If you shouldn't be able to copyright GRAPES...you shouldn't be able to copyright BOOKS.
Gemini should be trained on those books already, so in theory it could regurgitate some verbatim fragments (as NYT lawsuit against OpenAI showed some time ago).