AI Summary
Anthropic is rolling out a watermarking system for AI-generated text that embeds an invisible fingerprint detectable by machines, even surviving copy-pasting and light editing. The video explains the underlying algorithm, its limitations, and how to avoid it using open-weights models.
Chapters
Cloud AI (Anthropic) is watermarking generated text, rolling out now.
Text watermarking uses hidden characters or a fingerprint invisible to humans but detectable by machines, surviving copy-paste and some editing.
AI assigns probabilities to candidate words; a fingerprinting algorithm secretly colors words green (preferred) or red (not), nudging green words to appear more often.
Checking for watermark involves counting green words; many greens make human authorship extremely unlikely (probability less than winning the lottery).
Green words can be any word and change over time, making them hard to spot.
Likely uses SynthID variant with context-dependent probabilities and a tournament system, but core principle is same.
Watermark doesn't trace text to you, only shows Claude wrote or heavily edited it. Light editing doesn't remove it; full rewrite does.
Full rewrite or using an open-weights LLM can remove watermark.
Only eligible organizations can check for watermark, not individuals.
Use free and open-weights AI systems run locally to avoid watermarking.
AI text watermarking is a real, technically sophisticated measure, but it can be circumvented with full rewrites or open-weights models. The video advocates for open-source AI as a countermeasure.
Mentioned in this Video
💡 Key Takeaways
Invisible fingerprint
Explains a novel concept of text watermarking that is invisible to humans but detectable by machines.
00:26Detection method
Provides a clear, testable method for detecting watermarks based on word frequency.
01:39Removal methods
Offers practical advice on how to circumvent watermarking, which is crucial for users.
03:17Open-weights solution
Advocates for open-source AI as a countermeasure, aligning with the channel's philosophy.
03:44Full Transcript
[00:00] Images can be watermarked for copy protection. Now get this, Cloud AI announced that they are watermarking the text you generate with it. From when? When does this start?
[00:12] Well, Anthropic is rolling this out right now. Yup, I am not talking about this because I agree with it, but because I think it's important that all of you fellow scholars know about this, to inform the public.
[00:26] So this piece of text is watermarked. Wait, what? You can watermark an image by putting your logo on it, but text? How would you watermark text? You put hidden characters in it, right? Nope. This paper describes that it is a fingerprint
[00:43] in text that is invisible to humans, but is detectable for machines. It even survives copy-pasting and some editing too. I'll tell you what it doesn't survive in a minute.
[00:56] So when generating text the AI decides what the next word should be and there can be a few candidates Here you could say I saw a dog a puppy a cat or a house Based on context each of these words gets a probability
[01:13] to be chosen. Now, with a fingerprinting algorithm, it secretly assigns a color to each word. Some are green, preferred, some are red, not preferred. And now comes a little nepotism.
[01:27] A little cheating, if you will. When choosing the next word, the green ones get a little nudge upwards. They will occur slightly more often.
[01:39] So here's how to check for a watermark. In a piece of text, someone who knows the red and green words simply counts how many greens you have. This scheme has a mathematical property, where, as you see more and more green, the probability
[01:54] of it being real human text is extremely small. Found 21 grains in a paragraph suddenly the probability of that done by humans can be less than winning the lottery Much smaller Note that the green words can be anything no matter how inconspicuous so you can spot them Which words are green
[02:16] can also change over time. Ouch. This is the simplified version of the algorithm. They are likely using the SynthID variant, which has context-dependent probabilities,
[02:28] and a tournament system too. But the heart of the algorithm is the same in most research papers I read. Some words are preferred and are given a slight edge in the generation, creating a unique
[02:41] fingerprint. Now, there are a lot of misconceptions about this out there. One, Claude written text cannot be traced back to you, but it shows that Claude wrote the whole thing or heavily edited
[02:54] the text. I don't agree with this. I am making this video to let everyone know, knowledge only changes the world when it reaches people. Second, some say, just edit a few words, and it clean Nope You can get rid of it so easily So can you get rid of it Dear Fellow Scholars this is Two Minute Papers with Dr K Zsolnai With light editing no
[03:17] If you rewrite the whole thing, exchanging every word, yes, you can get rid of it. An open weights LLM that works for you can also help. Okay, so who can check if there is a watermark in the text?
[03:30] Well, not you and not me. Some eligible organizations can, but that's it for now. So what is the solution? Well, of course, use free and open weights AI systems and run them yourself.
[03:44] These work for you, not against you. That is the way of the scholar. We need new tools for the era of LLMs and Weights and Biases now has Weave, a lightweight
[03:56] toolkit to confidently iterate on LLM applications. Use traces to debug how data flows through each step of your app and use evaluations to measure your progress. It is the best! Try it out now at WNB.me slash papers,
[04:13] or click the link in the description below!