this is going to be the big thing in the next 100 years for me. where there is a human, there will be context, passion and meaning. it's the old saying: machine can tell you what, but only humans can tell you why.
How long before they change the terms and conditions to subtly claim ownership of your files? When you write code they already insert Co author attribution/
Say I write a text by hand And then I tell it to clean up the grammar and fix some sentences Did it make it?
They have said and tried some wild things already, like trying to get open weights effectively banned, which I believe they still think is in society's best interest (more that they think they know what's best for everyone)
The two competing legal arguments regarding copyright of LLM output are "it's like hiring a monkey" (author is the LLM, which is not a person, thus can't hold copyright and can't assign it to you) and "it's like taking a photograph" (the LLM is a machine through which the prompting person expresses their creativity, just like a camera). In no scenario is Anthropic the author of the work.
If we settle on the monkey analogy the Anthropic owns the monkey, but the owner of a monkey doesn't own copyright for the creations of the monkey. If we settle on the camera analogy, anthropic claiming ownership would be like Canon claiming they own pictures you take.
What Anthropic could do is to change the terms to give themselves a non-exclusive global perpetual license to use everything Claude makes
Watermarks should even be removable "if substantially proofread or altered" in the EU for the text from an LLM, that you verify to be as true as if you wrote it yourself.
Not clear is this is a may, shall, or should. It certainly is not a must.
Watermarks cane even be removed, "if substantially proofread" in the EU, law says. For the text from an LLM, that you verify to be as true as if you wrote it yourself.
Not clear is this is a may, shall, or should. It certainly is not a must.
If you post a comment here on HN, then hit the back button, edit your comment, and click “reply” again, you end up posting multiple comments. That’s what’s happened here.
The real problem is false positives. One false positives is enough to make the whole thing dangerous. The results can't really be acted upon without risking defamation. If you admit that you redistributed someone else's copyrighted work to an AI company that never forgets, it's an admission of distributing copyrighted works.
The law should have at the very least required offline validation tools that cannot track or retain a copy of the documents being checked.
This is really fascinating. Even the AI companies have incentives to reject AI generated content. It's like they want you to use AI for everything, but they don't want AI output fed back to them.
At work, right after an AI training, we were asked to use our "authentic" voice when writing mid year reviews.
Turbos don't inject exhaust air into an intake. Exhaust air spins the hot side of the turbo which causes the cool side to also spin. The cool side sucks in fresh air. Not exhaust air.
Your comment would make more sense if you mentioned exhaust gas recirculation (EGR) system.
Or maybe it's good for everybody to use LLMs for some things but not other things, as opposed to some people should use it for everything and others for nothing. These companies are definitely using LLMs for coding.
They still scrape code, I'd guess, e.g. from GitHub?
And there's tons of Claude-generated code there.
Also, I'd guess this is not just to prevent any AI-generated code in the training data, but specifically their own.
Percentage of users who put out their code on the web and also have a plan where Anthropic promises not to train on their data is problem also low.
So not excluding own code could be a real issue, since it would be impossible to deduplicate the training and RILHF data from their sessions with the code accessible elsewhere, and written by the very same users.
Could this be used to perform some sort of distillation or exploit? e.g. reminds me of the OWASP guideline on attack vectors where knowing if an ID is present or not in the database can be a form of exploit, like in password resets where they will say 'email foo@bar.com not found' rather than 'If foo@bar exists we have sent an email to foo@bar' or some other generic equivalent
Under what circumstances does the metadata get added? When Claude Code CLI needs a media file usually I just see it run an imagemagick or ffmpeg command to create one, which isn't going to have C2PA metadata.
If watermarking of any kind were actually going to prevent ai-generated “fraud” (or even ai being used as part of a scam), it would have more support on HN.
This kind of metadata injection increases surveillance without providing a meaningful deterrent.
Furthermore, it almost surely limits the quality of the output. At least the finterprinting of the text tokens has to, how can you otherwise constrain token output?
For example if you have 2 equally likely tokens to choose from you can pick the one specified by a key function instead of the one specified by an rng.
The actual implementation matters less than the fact that they can be perceived to be working on something.
You're never going to have a perfect solution here, and perfect is the enemy of good.
I'm wondering why they have restricted file types. You can't check a PDF for example... surely the main use case for people will be to check if a document was produced or edited by an LLM? That could be an attractive (if not misunderstood) proposition for academics
> I'm wondering why they have restricted file types. You can't check a PDF for example...
TFA/page actually seems incomplete. Text uses a completely different watermark format (an actual watermark as opposed to a provenance/authenticity signature), so it makes sense to me that they're not claiming to be able to scan PDFs when they can't yet incorporate that signal.
On a linked page, they say:
> Watermark detection is currently in private preview [...]
As far as I can tell, this and the recent change to add watermarking to text outputs[1], is to become compliant with the EU AI Act[2] and CA's AI Transparency Act[3], SB-924[4]. For large enough companies, all generated AI content is required to have watermarking.
Good point, re: the CA law. The rule for text output may be limited to the EU AI Act, which Anthropic notes as a reason in the linked article about its implementation.
Unfortunately this is just for Media? Some manual tells for Excel or PDFs is to check the author. Claude creates PDFs via wkhtmltopdf so the PDF Producer will be Qt and the Content Creator is wkhtmltopdf. Xlsx files are being created via openpyxl so in the metadata that is the author.
> Knowing where content came from, and whether AI was involved, makes it easier to trust what you see online.
That’s a cute way to imply their service is used to generate misinformation. They are basically saying to not trust the AI content made from their own product :)
Associations are generally two ways. You could possibly build something that just generated misinformation in a non-generalized way, but you cannot build a generalized writing machine that can't write misinformation.
I mean, I got what you were trying to say, I just turned it around on you using the same framework.
Like they should ban astronaut on a horse images or what? Misinformation becomes misinformation upon the fact of presenting it as so, not the fact of creation
What is interesting to me is that stripping the C2PA data is easy, but faking it is hard.
You can resave the file and the "made with Claude" signal disappears, but you cannot make a random file pass as Claude-made without Anthropic's signing key. So the useful guarantee is one-way. No signature means almost nothing.
"Faking" it is trivial. You don't need their signing keys when you can just ask them to sign whatever you like. Upload your own file with the prompt "present this file back to me again, as-is".
The goal of C2PA is that cameras will start to emit C2PA credentials. You will then have 3 situations:
* C2PA confirms a photo is authentic
* C2PA confirms a photo is AI generated
* C2PA missing, you don't know.
I reckon we will only see "C2PA missing" being treated as suspect in select situations (perhaps Reuters will require C2PA from their photojournalists, for example)
Camera C2PA can never meaningfully confirm that a photo is authentic, it bears about as much credence as EXIF metadata. It's like saying the existence of DRM confirms that a movie hasn't been pirated.
C2PA cryptographically guarantees that the bytes came from a hardware/software signer and that the signed payload has not been modified since that signature was applied.
So no, C2PA is not as easy to spoof as EXIF.
And no, the existence of DRM doesn't validate the integrity or the provenance of the bytes.
C2PA doesn't verify that you won the lotto, it only verifies that a Pixel camera captured that image and the pixels haven't been altered from what the camera captured.
Also, prisoners escaping from prison is not a theory either, prisoners have actually escaped. Does that mean all prisoners are free?
I assume the signing key is different from each camera unit(not only model), so if a picture of you winning lottery in US capture by a camera sold to someone in Thailand, it would be extreme unlikely to be real.
This is not the case with current systems, but could future systems be designed to be tamper-evident in a way that makes it impractical to sign fake images or extract the signing key without leaving evidence on the device?
If that were the case, I can imagine a subscription service in which you get a camera for some specified period of time, and then return it to the company that sold it for them to verify the camera hasn't been tampered with. Then the company could publish a list of which keys (unique per camera) have been verified to not be tampered with. Maybe this wouldn't stop everyone, but now the person trying to fake images has to re-do the process every so often and I imagine it's more expensive to avoid leaving evidence.
This might be too impractical to work, and it would be bad for privacy, but maybe for some people the tradeoffs actually would be worth it, someday. For now, I assume there are much cheaper and easier ways to detect faked images, at least for expert humans.
Your observation is sharp, but it's not just metainfo—it's load-bearing text. To remove it, you need to delve deep and alter the tapestry of carefully selected words.
> To detect watermarks embedded in text we have a Detection API which is currently in private preview to eligible organizations as required under EU law.
118 comments
[ 903 ms ] story [ 237 ms ] threadAny attempts they use are defeated by a text editor and CTRL SHIFT V. Unicode characters are no new thing.
Reminds me of how people tried to argue that NFTs aren't anything more than just jpegs.
Say I write a text by hand And then I tell it to clean up the grammar and fix some sentences Did it make it?
The two competing legal arguments regarding copyright of LLM output are "it's like hiring a monkey" (author is the LLM, which is not a person, thus can't hold copyright and can't assign it to you) and "it's like taking a photograph" (the LLM is a machine through which the prompting person expresses their creativity, just like a camera). In no scenario is Anthropic the author of the work.
If we settle on the monkey analogy the Anthropic owns the monkey, but the owner of a monkey doesn't own copyright for the creations of the monkey. If we settle on the camera analogy, anthropic claiming ownership would be like Canon claiming they own pictures you take.
What Anthropic could do is to change the terms to give themselves a non-exclusive global perpetual license to use everything Claude makes
C2PA is file metadata and can be trivially stripped away, unlike hidden watermarks like SynthID.
That would invalidate the hash on minor changes. Too much effort and not enough return.
For the text from an LLM, that you verify to be as true as if you wrote it yourself.
Not clear if making watermarks removable is a may, shall, or should, according to legislation. It certainly is not a must.
The law should have at the very least required offline validation tools that cannot track or retain a copy of the documents being checked.
As soon as the validator is available, just apply input fuzzing to defeat it.
At work, right after an AI training, we were asked to use our "authentic" voice when writing mid year reviews.
Your comment would make more sense if you mentioned exhaust gas recirculation (EGR) system.
Sorry.
And there's tons of Claude-generated code there.
Also, I'd guess this is not just to prevent any AI-generated code in the training data, but specifically their own.
Percentage of users who put out their code on the web and also have a plan where Anthropic promises not to train on their data is problem also low.
So not excluding own code could be a real issue, since it would be impossible to deduplicate the training and RILHF data from their sessions with the code accessible elsewhere, and written by the very same users.
This kind of metadata injection increases surveillance without providing a meaningful deterrent.
I then gave it the html as a file upload and it gave it back to me.
It's reasoning why? I already had the file so it must be mine. I told it a judge wouldn't care.
I'm wondering why they have restricted file types. You can't check a PDF for example... surely the main use case for people will be to check if a document was produced or edited by an LLM? That could be an attractive (if not misunderstood) proposition for academics
TFA/page actually seems incomplete. Text uses a completely different watermark format (an actual watermark as opposed to a provenance/authenticity signature), so it makes sense to me that they're not claiming to be able to scan PDFs when they can't yet incorporate that signal.
On a linked page, they say:
> Watermark detection is currently in private preview [...]
[1] https://www.anthropic.com/news/claude-text-watermark [2] https://digital-strategy.ec.europa.eu/en/policies/code-pract... [3] https://www.kqed.org/news/12095398/new-california-law-requir... [4] https://www.leginfo.legislature.ca.gov/faces/billTextClient....
That’s a cute way to imply their service is used to generate misinformation. They are basically saying to not trust the AI content made from their own product :)
I mean, I got what you were trying to say, I just turned it around on you using the same framework.
You can resave the file and the "made with Claude" signal disappears, but you cannot make a random file pass as Claude-made without Anthropic's signing key. So the useful guarantee is one-way. No signature means almost nothing.
* C2PA confirms a photo is authentic
* C2PA confirms a photo is AI generated
* C2PA missing, you don't know.
I reckon we will only see "C2PA missing" being treated as suspect in select situations (perhaps Reuters will require C2PA from their photojournalists, for example)
So no, C2PA is not as easy to spoof as EXIF.
And no, the existence of DRM doesn't validate the integrity or the provenance of the bytes.
It's like saying a prisoner has the same freedoms as everyone else because he could theoretically escape.
Here's a cryptographically signed + timestamped photo of me winning the lottery: https://verify.contentauthenticity.org/?source=https%3A%2F%2...
(Compare against winning numbers and draw timestamp at https://www.euro-millions.com/results/28-08-2026 )
Also, prisoners escaping from prison is not a theory either, prisoners have actually escaped. Does that mean all prisoners are free?
In cryptography, once any SINGLE person in the world has compromised a signing key, EVERY person in the world can use it.
Thus for a prisoner analogy, its equivalent to say once any SINGLE prisoner escapes, EVERY prisoner has the ability to escape.
So yes, in this terrible analogy it means all prisoners are free.
If that were the case, I can imagine a subscription service in which you get a camera for some specified period of time, and then return it to the company that sold it for them to verify the camera hasn't been tampered with. Then the company could publish a list of which keys (unique per camera) have been verified to not be tampered with. Maybe this wouldn't stop everyone, but now the person trying to fake images has to re-do the process every so often and I imagine it's more expensive to avoid leaving evidence.
This might be too impractical to work, and it would be bad for privacy, but maybe for some people the tradeoffs actually would be worth it, someday. For now, I assume there are much cheaper and easier ways to detect faked images, at least for expert humans.
1. Generate a bunch of responses with both Claude and various non-Claude LLMs (ChatGPT, Gemini, Kimi)
2. Train a discriminator model that can differentiate Claude vs. non-Claude
3. Train a de-watermarking model using the discriminator model as loss
edit: almost forgot the "—"
The honest seam is obvious: claudish works; imitation does not. That’s not nothing.
> To detect watermarks embedded in text we have a Detection API which is currently in private preview to eligible organizations as required under EU law.
https://support.claude.com/en/articles/16266773
https://www.anthropic.com/news/claude-text-watermark, Ctrl-F for "What about code?"
So… this utility appears to be pretty worthless.