6 comments of 10

[ 0.19 ms ] story [ 7.2 ms ] thread
In principle interesting, but I can't stand the Claude writing.

> Keys with real blast radius

> Here is what they unlock.

> This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used.

> We cloned the public dataset hub end to end: every repository, every branch, every large-file object

> The size is only half the story. These are the training sets behind models people actually use. The worst-hit ones are named, card-documented pretraining corpora that open models were built on. We verified every credential we cite against its provider, so they were live when we looked.

The whole post looks like a Claude artifact with random little cards.

It's also just too long, which is a side effect of using LLMs, it's just too easy to create walls of text.

(comment deleted)
Wouldn’t it just be easier to crawl the net? Not sure what huggingface has to do with anything here.
I think "7.6 PB" is more informative than "4.4x Empire State Building heights worth of DVDs", but that's just me.
This is fine. All of this is fine. :’)