In principle interesting, but I can't stand the Claude writing.
> Keys with real blast radius
> Here is what they unlock.
> This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used.
> We cloned the public dataset hub end to end: every repository, every branch, every large-file object
> The size is only half the story. These are the training sets behind models people actually use. The worst-hit ones are named, card-documented pretraining corpora that open models were built on. We verified every credential we cite against its provider, so they were live when we looked.
The whole post looks like a Claude artifact with random little cards.
It's also just too long, which is a side effect of using LLMs, it's just too easy to create walls of text.
I similarly find it interesting; I do not understand why it is not the default to just generate the thing as a draft, research anything you aren’t clear on, and re-write it in your own voice.
I guess Huggingface datasets are easy to crawl - those datasets are hosted to be crawled. In contrast to the web at large which will be behind Cloudflare etc
In principle interesting, but I can't stand the Claude writing.
> Keys with real blast radius
> Here is what they unlock.
> This is a floor, not an estimate of actual balances or unauthorized usage. The keys were verified but never used.
> We cloned the public dataset hub end to end: every repository, every branch, every large-file object
> The size is only half the story. These are the training sets behind models people actually use. The worst-hit ones are named, card-documented pretraining corpora that open models were built on. We verified every credential we cite against its provider, so they were live when we looked.
The whole post looks like a Claude artifact with random little cards.
It's also just too long, which is a side effect of using LLMs, it's just too easy to create walls of text.
Ever read a book by Yuval Noah Harari? He seems to have been writing like this since before LLMs. Perhaps that’s where they got it from.
I similarly find it interesting; I do not understand why it is not the default to just generate the thing as a draft, research anything you aren’t clear on, and re-write it in your own voice.
That takes a level of taste and craft that many in this industry do not possess.
Wouldn’t it just be easier to crawl the net? Not sure what huggingface has to do with anything here.
I guess Huggingface datasets are easy to crawl - those datasets are hosted to be crawled. In contrast to the web at large which will be behind Cloudflare etc