Yeah but it's like saying "Masterlocks will always be easy to pop off with a hammer". Of course, but by doing so you are actively engaging in fraud, which then puts the onus on you and whoever you are attempting to deceive.
I feel like this is going to end up being like cookie laws. It sounds good, I don't know how any one benefits from it.
Currently I can recognize AI text because I read thousands of ai generated text. I know that 110% of yahoo finance news is generated. I don't want to read an AI generated personal blog, but if I do what's the problem really? Other than the companies distinguishing AI text for getting better training data, how do people benefit from watermarked text exactly?
>I feel like this is going to end up being like cookie laws. It sounds good, I don't know how any one benefits from it.
Does it sound good?
Either anthropic does not tell anyone the signal and only they can verify the text. Or they share the signal and everyone can verify it, which means everyone can bypass it and its just an inconvenience.
> I feel like this is going to end up being like cookie laws
Companies will implement it in the worst way possible? Nowhere in the cookie laws does it say you need to add a banner, you just can't spy on users without their consent.
I would income education institutions might benefit from being able to distinguish between AI and human generated text. The people who benefit, in the long run, are the students.
The article mentions you can always simply use a smaller, local, un-watermarked LLM to rephrase the original watermarked text. Which is true, sure.
But if we're talking about deterministically taking some watermarked LLM output and having a function removeWatermark(text), it won't necessarily be "trivial" to remove, because the watermark function itself need not be public. Only the API that tests for the watermark need be public, right?
Anthropic's magic watermark could be, like the article mentions, something like "every 7th semicolon has a N% chance to be a comma where N is the sum of the last X characters mod Y, and every character in the bit range q1...q2 has a Z% chance to..." etc etc etc. And if Anthropic controls those variables, it would be very difficult to determine the rule, even with some pretty advanced analysis (I would assume). And keep in mind, that example rule I mentioned is pretty naive, too. I expect the actual rule would be way more advanced and not so straightforward as "swap every <charX> for a <charY>"
>Anthropic's magic watermark could be, like the article mentions, something like "every 7th semicolon has a N% chance to be a comma where N is the sum of the last X characters mod Y, and every character in the bit range q1...q2 has a Z% chance to..." etc etc etc.
Doesnt this and probably all techniques require the validator to know which portion of the text to validate?
If its not all generated together then how could it reliably carry the mark? Sure, run it against the full text. But what if the full text was not one-shot by the llm?
In other words, in order to reliably detect if the text is ai you need to first determine which part of the text was generated together by ai.
To remove the watermark you'd just need to paraphrase the text with another model that wasn't adding the statistical signature to it. You wouldn't need to know what the original statistical signature was.
"What would it even look like to sign ChatGPT outputs? There’s no artifact to pass around."
It would look like a lot of little signatures on little bits of text, and then larger signatures on a collection of those chunks once the larger chunk exists. It's not that hard. It's just a lot of signatures.
But this scheme could only ever prove that this bit of text was made by a given AI, and validate anything else ever included in the signature hasn't been tampered with. It's not hard to work up a scheme that proves (within reason) a text was generated no earlier than some date by incorporating some sort of information that could only have been known at that date so that could be validated. But this isn't even a step in the direction of proving that something was made by a human. And that's assuming the private keys stay private, which is its own tricky problem. If a private key ever leaks anything signed with it becomes invalidated.
For SynthID and similar solutions, there is much I don't understand ...
Here's what I grasp: The AI system scores each token and then selects tokens based on those scores. If we encode something in the token selection routine ('in order choose the 1st, 3rd, 1st, 5th, 2nd, then 1st highest scored tokens'), we can identify AI-generated text by comparing sample text (ST) to the expected text (ET) for that prompt.
1) How do we score the tokens for the ET without the original prompt? Even a Markov-like process needs to start somewhere.
2) To recreate ET don't we need to maintain, until the end of time, the AI state - entire model and code - at the time of ST output?
3) Doesn't #2 require maintaining all states for all AIs? Often you won't know when and from which AI system the ST might have been generated. What happens when an AI vendor goes out of business?
4) To recreate ET, don't we effectively have to rerun the prompt? Won't rerunning it for every verification increase most costs of AI output by an order of magnitude? Most of what AI vendors do would be ST validation.
Society can overcome this problem by changing the way we think about text. Raw plain text should be banned, all text is cryptographically signed by the editor, gui element, or tool that created it.
The goal of the AI act is not to determine if an "oh yeah!" comment was AI generated. The target is long papers that falsely claim human review and can have real significant consequences.
E.g. research paper, law makers, lawyers, state policies, notaries,...
These are much longer content and thus statistically they will disclose a better guess at AI generated content.
Asking another AI to paraphrase will not erase the mark (which they are unaware about) but rather cumulatively add their own mark and make it easier to detect.
The problem is not to use AI, but to endorse the responsibility of the content you (as a human) deliver and somehow make sure that fake-news, biased content or unverified output is detected as early as possible.
Just look at global warming, caused by our totally negligent behaviour.... AI will decimate people's thought processes and they essentially become zombie like. And it will take a few decades until they realize what have they done. It will be a painful experience to shoo off people from the heroin called AI in order for them to use their brains again.
AI watermarks feel like they're approaching the problem from the wrong side - no matter what it will be possible to remove the watermark (Via manual rewriting, local LLLMs, etc). Instead it seems like we need "proof of human creation". And the only way I see that being possible is hardware-attested proof of keypresses. Which obviously has huge privacy implications, but how else would you actually know that a piece of content was produced by a human pressing keys on a keyboard? The proof would also need to include timestamps for the keypresses, so tell if someone is just copying from another window.
That sounds like an interesting idea, but I think there's always weird edge cases. Like what if I pressed the keys to write an essay that was actually generated by an LLM. It would be hardware-attested, but I just transcribed most of my sentences from an LLM and maybe I even introduce my own typos or sentences because I am half-assing the transcription.
You could do that, but a real human doesn't just type things out in one straight shot (modulo typos that are fixed "in-line"). There's going to be editing, pauses, copy-pasting, moving things around, etc. And at worst it puts a cap on how much text you can generate since at minimum it does need to be manually retyped.
Whatever happened to just delivering the best product or service? Why must tech be full of ninnying nannies that act against their users, "for their 'safety'‽"
Because the best product should still take care of some social accountability, Meta would not sell their shitty glasses without the light indicator and even with it, as it's trivial to hack there's large social pushback.
It is not the product's job or role to make the customer act in a socially pleasing manner. It seems to happen because a customer-sovereignty respecting competitor hasn't come in to clean up the moral crusader's lunch.
Every technology, every action has a moral aspect to it.
Just because you want to ignore it doesn't mean it's not there.
Society then decides on what morals we want to uphold, or at least tolerate, vs. those we want to change.
Laws are hard power to change behavior. Influence/goodwill/etc. are soft power ways to change behavior and bend organizations in ways they may not otherwise want to.
Would you pay a pto keep grandma from eating dogfood if it wasn't the law that you have to put a portion of your earnings into Medicare and Social Security?
I don't think this is a good argument. There exist digital watermarking techniques for image and video that are imperceptible to humans but still survive cropping, rotation, resizing, recompression, or an analog round trip (photographing the image or pointing a camera at the video). These watermarks aren't, like, hiding in the low-bits of color information, they're spread among many perceptible details.
It's not clear to me that it's impossible, or even especially difficult, to make something that survives a casual LLM paraphrase. Remember, all you need to encode is a single bit of info. There's a lot of space to redundantly encode that signal.
“Everyone will simply do fraud” - it is a bizarrely immoral world that the LLMs seem to have unleashed on us. Tech was kinda heading that way anyway, but our friends the magic robots really seem to have turbocharged it.
Maybe this is an information theory thing but can there exist a n“ I am human” shibboleth or is this concept fundamentally impossible?
It feel that today you can generally convince someone with text alone that you’re human. Beepity boopity zip zap zoopity today’s AIs aren’t this loose and derpy. Here’s a fTypo and my secret stash of dashes ——-–.
Isnt this absurd? Say im brainstorming a resume bulet point. Its 15 words. I like it but want to condense it to a single line. I give it to the ai and tell it how much overflows and now it gives me back a simplified sentence and 11 words and some extra stenography constraint? What kind of rule could possibly not effect the quality of that output?
Ok say i do that on 30% of bullet points. Karen the hiring manager is vehemently anti AI. She gives my resume to her AI scanner and what does she find? This not a rhetorical question. Will it treat the text as a whole and not find it? Does it scan every combination of contiguous terms? It could scan bullet points but i could generate in pairs of 2. What about novels?
https://declaude.org/watermarking/ did a good job in explaining how SynthID works. As per their blog, it feels like it will be difficult to remove watermarking on bigger text and the checking for watermarking is also not complex
I've been working in the media and model IP space for quite many years. What this article misses a bit is the threat model for watermarking in general.
Watermarking and fingerprinting have inherently weak security guarantees -- they rely a lot on security through obscurity, weak assumed adversaries to deliver.
There are clear trade offs between true positives, false positives and maintaining the quality of the media. It's true for audio-visual media, models and their outputs alike.
As much as I like to take shots at poor technical choices by corps and govs, this one is unjustified. Sure, inserting glyphs is bad but biased sampling is as good as it gets in 2026.
Even if you find a way to 100% watermark any text, couldn't you just use a non-watermark model? I have a hard time believing every AI company on heart would comply.
Hell, even if every AI company on earth decide to somehow apply watermarking to their next model, they would also need to apply it to all the previous version that they commercialize. Given that anybody could make a copy of an open source model right now and would be safe forever, this seems quite the lost cause to me.
46 comments
[ 55.3 ms ] story [ 1095 ms ] threadjust make an API that returns the string distance between a previously generated paragraph and the query?
that would sidestep this whole problem class.
regulators could even specify how that has to work.
what am i missing?
People underestimate the value of rules that only take malice and a little knowledge to break.
And they tend to exaggerate that underestimation if they... don't like the rule.
Currently I can recognize AI text because I read thousands of ai generated text. I know that 110% of yahoo finance news is generated. I don't want to read an AI generated personal blog, but if I do what's the problem really? Other than the companies distinguishing AI text for getting better training data, how do people benefit from watermarked text exactly?
Does it sound good?
Either anthropic does not tell anyone the signal and only they can verify the text. Or they share the signal and everyone can verify it, which means everyone can bypass it and its just an inconvenience.
Companies will implement it in the worst way possible? Nowhere in the cookie laws does it say you need to add a banner, you just can't spy on users without their consent.
This will most likely bring a sift end to at least the low hanging fruit.
But if we're talking about deterministically taking some watermarked LLM output and having a function removeWatermark(text), it won't necessarily be "trivial" to remove, because the watermark function itself need not be public. Only the API that tests for the watermark need be public, right?
Anthropic's magic watermark could be, like the article mentions, something like "every 7th semicolon has a N% chance to be a comma where N is the sum of the last X characters mod Y, and every character in the bit range q1...q2 has a Z% chance to..." etc etc etc. And if Anthropic controls those variables, it would be very difficult to determine the rule, even with some pretty advanced analysis (I would assume). And keep in mind, that example rule I mentioned is pretty naive, too. I expect the actual rule would be way more advanced and not so straightforward as "swap every <charX> for a <charY>"
Doesnt this and probably all techniques require the validator to know which portion of the text to validate?
If its not all generated together then how could it reliably carry the mark? Sure, run it against the full text. But what if the full text was not one-shot by the llm?
In other words, in order to reliably detect if the text is ai you need to first determine which part of the text was generated together by ai.
It would look like a lot of little signatures on little bits of text, and then larger signatures on a collection of those chunks once the larger chunk exists. It's not that hard. It's just a lot of signatures.
But this scheme could only ever prove that this bit of text was made by a given AI, and validate anything else ever included in the signature hasn't been tampered with. It's not hard to work up a scheme that proves (within reason) a text was generated no earlier than some date by incorporating some sort of information that could only have been known at that date so that could be validated. But this isn't even a step in the direction of proving that something was made by a human. And that's assuming the private keys stay private, which is its own tricky problem. If a private key ever leaks anything signed with it becomes invalidated.
Here's what I grasp: The AI system scores each token and then selects tokens based on those scores. If we encode something in the token selection routine ('in order choose the 1st, 3rd, 1st, 5th, 2nd, then 1st highest scored tokens'), we can identify AI-generated text by comparing sample text (ST) to the expected text (ET) for that prompt.
1) How do we score the tokens for the ET without the original prompt? Even a Markov-like process needs to start somewhere.
2) To recreate ET don't we need to maintain, until the end of time, the AI state - entire model and code - at the time of ST output?
3) Doesn't #2 require maintaining all states for all AIs? Often you won't know when and from which AI system the ST might have been generated. What happens when an AI vendor goes out of business?
4) To recreate ET, don't we effectively have to rerun the prompt? Won't rerunning it for every verification increase most costs of AI output by an order of magnitude? Most of what AI vendors do would be ST validation.
I was assuming it was something like SynthID rather than just sneaky invisible unicode but it's hard to tell from the description.
E.g. research paper, law makers, lawyers, state policies, notaries,...
These are much longer content and thus statistically they will disclose a better guess at AI generated content.
Asking another AI to paraphrase will not erase the mark (which they are unaware about) but rather cumulatively add their own mark and make it easier to detect.
The problem is not to use AI, but to endorse the responsibility of the content you (as a human) deliver and somehow make sure that fake-news, biased content or unverified output is detected as early as possible.
Whatever happened to just delivering the best product or service? Why must tech be full of ninnying nannies that act against their users, "for their 'safety'‽"
Just because you want to ignore it doesn't mean it's not there.
Society then decides on what morals we want to uphold, or at least tolerate, vs. those we want to change.
Laws are hard power to change behavior. Influence/goodwill/etc. are soft power ways to change behavior and bend organizations in ways they may not otherwise want to.
Would you pay a pto keep grandma from eating dogfood if it wasn't the law that you have to put a portion of your earnings into Medicare and Social Security?
It's not clear to me that it's impossible, or even especially difficult, to make something that survives a casual LLM paraphrase. Remember, all you need to encode is a single bit of info. There's a lot of space to redundantly encode that signal.
It feel that today you can generally convince someone with text alone that you’re human. Beepity boopity zip zap zoopity today’s AIs aren’t this loose and derpy. Here’s a fTypo and my secret stash of dashes ——-–.
But one day even this won’t do, right?
Ok say i do that on 30% of bullet points. Karen the hiring manager is vehemently anti AI. She gives my resume to her AI scanner and what does she find? This not a rhetorical question. Will it treat the text as a whole and not find it? Does it scan every combination of contiguous terms? It could scan bullet points but i could generate in pairs of 2. What about novels?
Watermarking and fingerprinting have inherently weak security guarantees -- they rely a lot on security through obscurity, weak assumed adversaries to deliver.
There are clear trade offs between true positives, false positives and maintaining the quality of the media. It's true for audio-visual media, models and their outputs alike.
As much as I like to take shots at poor technical choices by corps and govs, this one is unjustified. Sure, inserting glyphs is bad but biased sampling is as good as it gets in 2026.
EDIT: grammar
Even if you find a way to 100% watermark any text, couldn't you just use a non-watermark model? I have a hard time believing every AI company on heart would comply.
Hell, even if every AI company on earth decide to somehow apply watermarking to their next model, they would also need to apply it to all the previous version that they commercialize. Given that anybody could make a copy of an open source model right now and would be safe forever, this seems quite the lost cause to me.