48 comments

[ 0.29 ms ] story [ 19.5 ms ] thread
(comment deleted)
(comment deleted)
I was not asked for permission when they trained their model on my output.
I added © 2024 - No Rights for Hypocrites to the footer of my website. Claude scraped all pages at least once every week since it came online.

Now what.

You can, they are just confused about what effective altruism means.
Corollary: If you can't use the output to train, then you don't own them.
They stole all the data, and then dont want you to steal it back. Its basically Robin Hood all over again.
Honestly, the entire dev community should save question and answers, uploaded them to a shared repo anonymously, then just use that data to distill further models and provide them to the public for free.

Theoretically, this should be legal and ethical, when comparing to Anthropic's own behavior. That said, the reason you can't is Anthropic states in their terms that they don't want you to do this.

Anthropic's entire business model is skating on thin ice.

This seems like the standard AI company hypocrisy. Hopefully Claude users disregard this nonsense.
Yeah, no. I was OK with them scraping everything if it means we get AI, but, conversely, they don't get to control what happens to their outputs.

Hell, arguably they should release their weights (or at least the weights of their older models), since they trained them on the concentrated knowledge of humankind.

If more and more of the web's content is AI generated, AI companies are bound to train on each other's data.

Or, what if I generate content with Claude/ChatGPT/Gemini, warp it in HTML using an open model, put this on my website conveniently dedicated to "Best practices in prompt and AI answers" for example, then train my own model that only scraps my website?

You downloaded millions of books from shadow libraries and trained on them.

I don't care about your policies, go fuck yourself.

> When customers use Claude to generate Outputs that then train competing models, they're essentially using our infrastructure and investment to build direct competitors to our service

We did so, please do not repeat it at home.

From my interpretation, in order to get the data you 'own' (it's not theirs to give away since they can't claim the copyright on it), you need to use their services. The agreement the user has with Antropic is for the service, not a restriction on how the data is used.
So can you write GPL content using Claude? MIT? Because if this condition applies to the output, I don't see how it's compatible with FOSS.
Now that is standard AI company hypocrisy
Imagine that you bought an axe, but the manufacturer banned you from making more axes with it.
All this posturing is all for naught, because people can do proper logit distillation into smaller models with open-weight models.
There must be a clear difference between terms of the Anthropic service and the legal standing of the ai output. The output is mine and I'll do with it whatever I want. The service is Anthropic's and they can do business with whoever they want. Everything else is hallucination.
Wouldn't such language and reasoning from Anthropic be an argument that they needed written permission to train their model on data from websites?

Has any individual somewhere around the world tested this in court by now? Sued Anthropic for copyright infringement because Claude can reproduce information that is only available on their website?

It shouldn't be that expensive, right? If you sue them for - say - $10000 then what would the costs of such a court case be?

Personally, I think "learning" is not a copyright violation. But if they themselves make it one, then they should also face the consequences, no?

But you can give them all to me and I can train my model on them. The restriction on you is not related to your ownership of that days, it's related to the contract you "signed".

Or you can publish them on the web and countless others will do it.