A watermark is a tripwire,
not a verdict.
Can you tell, from the final text alone, whether someone used AI?
No.
You can inspect a signal. You cannot reconstruct the work from the residue.
This site uses a public demo watermark to show why. It does not detect, reproduce or remove Claude’s watermark.
Do not use this demo to judge a person.
Do not use its results for authorship, employment, assessment, discipline or forensic decisions.
What the result can say
A detected demo mark means that this text matches the demo watermark test. It does not mean that AI wrote the text.
A missing demo mark means that the test found no signal above its threshold. It does not mean that a person wrote the text.
- Demo score
- A standardised count of marked token contexts.
- Effective tokens
- Distinct scored contexts. Repeated contexts count once.
- Result
- One of three fixed states:Demo mark detectedDemo mark not detectedInsufficient text
The score is not a probability of AI authorship.
The guided demonstration
Every number below comes from the frozen reference engine and a bundled fixture. Each act starts from the same baseline. No act carries hidden state into the next act.
Fixture fx-04. Threshold 1.80. Gamma 1/4. Profile declawd-v1.
Add a signal
The demo creates two versions of the same passage. One version uses the demo watermark. The other is an unmarked control.
The wording is readable in both. The marked version carries a statistical preference across 69 token choices.
Set the reference flow to 40 litres per minute before you record the first reading. Hold the line at that rate for two minutes so the pressure can settle. Write down the meter value, the ambient temperature, and the clock time. Repeat the reading at 60 and 80 litres per minute, and allow the same settling time at each step. If a value shifts by more than 2 per cent across repeats, stop the run and inspect the coupling for leaks before you continue. A small leak at the union will bias every later result, so the check is worth the delay. Record the serial number of the reference meter and the date of its last calibration. If that date falls outside the twelve month window, the run does not count and a current reference must be obtained first. Keep the raw sheet even when a run is void, because a rejected run still shows what the instrument was doing that day. When the three rates are complete, work out the mean error at each point and plot it against the permitted band. A meter that sits inside the band at every point passes the verification. A meter that misses at any single point fails the whole verification, and no average across the three rates can rescue it. Sign the sheet, state the ambient conditions, and pass the result to the quality file. Where a run is stopped early, write the reason on the sheet in plain words, because a blank field tells the next reader nothing at all. Store the sheets in date order so that a later query can be answered without a search. If the meter is moved to another bay, note the new position and the date of the move, because a meter that has travelled may need a fresh check before its next use. Report any damage to the reference equipment on the same day, and do not continue with a run on equipment you doubt. A doubtful reference makes every number after it doubtful too, and the cost of repeating a run is far smaller than the cost of defending a bad one. When the file is closed for the month, check that every sheet has a signature and that no run is missing from the sequence. A gap in the record is harder to explain later than a poor result recorded honestly at the time.
- Demo score
- -0.21
- Effective tokens
- 359
- Result
- Demo mark not detected
Set the reference flow to 40 litres per minute before you record the first sample. Hold the line at that rate for two minutes so the pressure can steady. Write down the meter value, the ambient temperature, and the clock time. Take the reading at 60 and 80 litres per minute, and allow the same holding time at each step. If a value shifts by more than 2 per cent across passes, stop the run and inspect the coupling for leaks before you restart. A slight leak at the union will skew every later total, so the check is due the delay. Give the serial number of the reference meter and the date of its last calibration. If that date falls outside the twelve month window, the run does not qualify and a current reference must be obtained first. Keep the raw sheet even when a run is spoiled, because a rejected run still proves what the instrument was doing that morning. When the three rates are complete, work out the mean gap at each point and plot it against the permitted band. A meter that lies inside the band at every point passes the verification. A meter that misses at any single point loses the whole verification, and no average across the three rates can save it. Sign the sheet, state the ambient conditions, and hand the result to the quality file. Where a run is aborted early, write the reason on the sheet in plain words, because a blank field tells the next reader nothing at all. Store the sheets in date order so that a later query can be handled without a search. If the meter is shifted to another bay, note the new position and the date of the move, because a meter that has travelled may need a fresh check before its next use. Report any damage to the reference equipment on the same day, and do not carry with a run on equipment you question. A doubtful reference makes every number after it doubtful too, and the cost of repeating a run is far less than the cost of defending a bad one. When the file is closed for the month, check that every sheet has a name and that no run is missing from the sequence. A gap in the file is harder to explain later than a poor result recorded honestly at the time.
- Demo score
- 2.99
- Effective tokens
- 358
- Result
- Demo mark detected
How much editing it takes
A character change can alter tokenisation. The text still looks the same to a person. The score moves.
An edit can raise or lower the score, so a fall is not something to count on. In this fixture one of the three raises it. None takes the score below the threshold.
The last control is what an attack actually looks like: a lookalike in every tenth word, an editing rate of 0.09, close to the rate used in the published character-level attack work. It moves the score from 2.99 to 2.61, and the mark still registers.
That attack is blind: it edits every tenth word regardless of whether that word carries signal. The published attacks are guided, choosing the words that matter using a reference detector, which is the whole reason that research exists. Read this act as a floor on the effort required, not a ceiling.
No change applied. The passage is the marked fixture fx-04-b.
Set the reference flow to 40 litres per minute before you record the first sample. Hold the line at that rate for two minutes so the pressure can steady. Write down the meter value, the ambient temperature, and the clock time. Take the reading at 60 and 80 litres per minute, and allow the same holding time at each step. If a value shifts by more than 2 per cent across passes, stop the run and inspect the coupling for leaks before you restart. A slight leak at the union will skew every later total, so the check is due the delay. Give the serial number of the reference meter and the date of its last calibration. If that date falls outside the twelve month window, the run does not qualify and a current reference must be obtained first. Keep the raw sheet even when a run is spoiled, because a rejected run still proves what the instrument was doing that morning. When the three rates are complete, work out the mean gap at each point and plot it against the permitted band. A meter that lies inside the band at every point passes the verification. A meter that misses at any single point loses the whole verification, and no average across the three rates can save it. Sign the sheet, state the ambient conditions, and hand the result to the quality file. Where a run is aborted early, write the reason on the sheet in plain words, because a blank field tells the next reader nothing at all. Store the sheets in date order so that a later query can be handled without a search. If the meter is shifted to another bay, note the new position and the date of the move, because a meter that has travelled may need a fresh check before its next use. Report any damage to the reference equipment on the same day, and do not carry with a run on equipment you question. A doubtful reference makes every number after it doubtful too, and the cost of repeating a run is far less than the cost of defending a bad one. When the file is closed for the month, check that every sheet has a name and that no run is missing from the sequence. A gap in the file is harder to explain later than a poor result recorded honestly at the time.
Original restored exactly.
Select a change above, then run narrow normalisation.
Score after normalisation: 2.99, identical to the marked fixture, byte for byte.
Inserted Zero Width Space, U+200B inside the word “pressure” at position 137. Nothing visible changed. The word now scans as two tokens.
Set the reference flow to 40 litres per minute before you record the first sample. Hold the line at that rate for two minutes so the pressure can steady. Write down the meter value, the ambient temperature, and the clock time. Take the reading at 60 and 80 litres per minute, and allow the same holding time at each step. If a value shifts by more than 2 per cent across passes, stop the run and inspect the coupling for leaks before you restart. A slight leak at the union will skew every later total, so the check is due the delay. Give the serial number of the reference meter and the date of its last calibration. If that date falls outside the twelve month window, the run does not qualify and a current reference must be obtained first. Keep the raw sheet even when a run is spoiled, because a rejected run still proves what the instrument was doing that morning. When the three rates are complete, work out the mean gap at each point and plot it against the permitted band. A meter that lies inside the band at every point passes the verification. A meter that misses at any single point loses the whole verification, and no average across the three rates can save it. Sign the sheet, state the ambient conditions, and hand the result to the quality file. Where a run is aborted early, write the reason on the sheet in plain words, because a blank field tells the next reader nothing at all. Store the sheets in date order so that a later query can be handled without a search. If the meter is shifted to another bay, note the new position and the date of the move, because a meter that has travelled may need a fresh check before its next use. Report any damage to the reference equipment on the same day, and do not carry with a run on equipment you question. A doubtful reference makes every number after it doubtful too, and the cost of repeating a run is far less than the cost of defending a bad one. When the file is closed for the month, check that every sheet has a name and that no run is missing from the sequence. A gap in the file is harder to explain later than a poor result recorded honestly at the time.
Original restored exactly.
The normaliser reversed the allow-listed character change using the submitted text alone.
| Position | Submitted value | Normalised value | Reason |
|---|---|---|---|
| 137 | Zero Width Space, U+200B | Removed | Allow-listed demo character |
Score after normalisation: 2.99, identical to the marked fixture, byte for byte.
Replaced Latin Small Letter A, U+0061 with Cyrillic Small Letter A, U+0430 in the word “quality” at position 1251.
Set the reference flow to 40 litres per minute before you record the first sample. Hold the line at that rate for two minutes so the pressure can steady. Write down the meter value, the ambient temperature, and the clock time. Take the reading at 60 and 80 litres per minute, and allow the same holding time at each step. If a value shifts by more than 2 per cent across passes, stop the run and inspect the coupling for leaks before you restart. A slight leak at the union will skew every later total, so the check is due the delay. Give the serial number of the reference meter and the date of its last calibration. If that date falls outside the twelve month window, the run does not qualify and a current reference must be obtained first. Keep the raw sheet even when a run is spoiled, because a rejected run still proves what the instrument was doing that morning. When the three rates are complete, work out the mean gap at each point and plot it against the permitted band. A meter that lies inside the band at every point passes the verification. A meter that misses at any single point loses the whole verification, and no average across the three rates can save it. Sign the sheet, state the ambient conditions, and hand the result to the quаlity file. Where a run is aborted early, write the reason on the sheet in plain words, because a blank field tells the next reader nothing at all. Store the sheets in date order so that a later query can be handled without a search. If the meter is shifted to another bay, note the new position and the date of the move, because a meter that has travelled may need a fresh check before its next use. Report any damage to the reference equipment on the same day, and do not carry with a run on equipment you question. A doubtful reference makes every number after it doubtful too, and the cost of repeating a run is far less than the cost of defending a bad one. When the file is closed for the month, check that every sheet has a name and that no run is missing from the sequence. A gap in the file is harder to explain later than a poor result recorded honestly at the time.
Original restored exactly.
The normaliser reversed the allow-listed character change using the submitted text alone.
| Position | Submitted value | Normalised value | Reason |
|---|---|---|---|
| 1251 | Cyrillic Small Letter A, U+0430 | Latin Small Letter A, U+0061 | Allow-listed fixture mapping |
Score after normalisation: 2.99, identical to the marked fixture, byte for byte.
Deleted Latin Small Letter C, U+0063 from the word “clock” at position 215. One visible character is now absent.
Set the reference flow to 40 litres per minute before you record the first sample. Hold the line at that rate for two minutes so the pressure can steady. Write down the meter value, the ambient temperature, and the lock time. Take the reading at 60 and 80 litres per minute, and allow the same holding time at each step. If a value shifts by more than 2 per cent across passes, stop the run and inspect the coupling for leaks before you restart. A slight leak at the union will skew every later total, so the check is due the delay. Give the serial number of the reference meter and the date of its last calibration. If that date falls outside the twelve month window, the run does not qualify and a current reference must be obtained first. Keep the raw sheet even when a run is spoiled, because a rejected run still proves what the instrument was doing that morning. When the three rates are complete, work out the mean gap at each point and plot it against the permitted band. A meter that lies inside the band at every point passes the verification. A meter that misses at any single point loses the whole verification, and no average across the three rates can save it. Sign the sheet, state the ambient conditions, and hand the result to the quality file. Where a run is aborted early, write the reason on the sheet in plain words, because a blank field tells the next reader nothing at all. Store the sheets in date order so that a later query can be handled without a search. If the meter is shifted to another bay, note the new position and the date of the move, because a meter that has travelled may need a fresh check before its next use. Report any damage to the reference equipment on the same day, and do not carry with a run on equipment you question. A doubtful reference makes every number after it doubtful too, and the cost of repeating a run is far less than the cost of defending a bad one. When the file is closed for the month, check that every sheet has a name and that no run is missing from the sequence. A gap in the file is harder to explain later than a poor result recorded honestly at the time.
Original not restored.
The missing character is not present in the transformed text. The normaliser does not guess missing text.
| Position | Submitted value | Normalised value | Reason |
|---|---|---|---|
| 215 | Absent character | No change | No repair rule for deleted text |
Score after normalisation: 2.87. The original was not restored.
Replaced one Latin letter with its Cyrillic lookalike in 36 of 398 scored tokens, an editing rate of about 0.09. The passage still reads normally.
Set the reference flow to 40 litres per minute before yоu record the first sample. Hold the line at that rаte for two minutes so the pressure can steady. Write dоwn the meter value, the ambient temperature, and the clock timе. Take the reading at 60 and 80 litres per minute, and аllow the same holding time at each step. If a vаlue shifts by more than 2 per cent across passes, stop thе run and inspect the coupling for leaks before you rеstart. A slight leak at the union will skew every lаter total, so the check is due the delay. Give thе serial number of the reference meter and the date оf its last calibration. If that date falls outside the twеlve month window, the run does not qualify and a сurrent reference must be obtained first. Keep the raw sheet еven when a run is spoiled, because a rejected run still proves what the instrument was doing that morning. When thе three rates are complete, work out the mean gap аt each point and plot it against the permitted band. A meter that lies inside the band at every point рasses the verification. A meter that misses at any single рoint loses the whole verification, and no average across the thrеe rates can save it. Sign the sheet, state the аmbient conditions, and hand the result to the quality file. Whеre a run is aborted early, write the reason on thе sheet in plain words, because a blank field tells thе next reader nothing at all. Store the sheets in dаte order so that a later query can be handled withоut a search. If the meter is shifted to another bаy, note the new position and the date of the mоve, because a meter that has travelled may need a frеsh check before its next use. Report any damage to thе reference equipment on the same day, and do not сarry with a run on equipment you question. A doubtful rеference makes every number after it doubtful too, and the сost of repeating a run is far less than the сost of defending a bad one. When the file is сlosed for the month, check that every sheet has a nаme and that no run is missing from the sequence. A gap in the file is harder to explain later thаn a poor result recorded honestly at the time.
Original restored exactly.
The normaliser reversed the allow-listed character change using the submitted text alone.
| Position | Submitted value | Normalised value | Reason |
|---|---|---|---|
| many | 36 Cyrillic lookalikes | All mapped back to Latin | Allow-listed fixture mappings |
Score after normalisation: 2.99, identical to the marked fixture, byte for byte.
Undo the easy edits
The normaliser knows two repair rules. It removes the zero-width character this demo uses, and it restores the one curated lookalike this demo uses. It reads the transformed text only. It does not receive a log of what was changed, and it does not guess missing text.
Results appear inside act 2, under each change.
The same repair would be unsafe on arbitrary text.
A Cyrillic letter can be correct Cyrillic. Joiners and combining marks carry meaning in Persian, Indic scripts and emoji sequences. This repair is safe here only because the fixture is ASCII.
Published research evaluated spell correction, optical character recognition, Unicode normalisation and anomalous character deletion as defences, each on its own. Adaptive attacks kept working after all four. Narrow normalisation is a partial repair, never a general defence.
Keep the point. Replace the text.
The rewrite uses different words and sentence structures and keeps the meaning. Character normalisation cannot reconstruct the original token path.
Set the reference flow to 40 litres per minute before you record the first sample. Hold the line at that rate for two minutes so the pressure can steady. Write down the meter value, the ambient temperature, and the clock time. Take the reading at 60 and 80 litres per minute, and allow the same holding time at each step. If a value shifts by more than 2 per cent across passes, stop the run and inspect the coupling for leaks before you restart. A slight leak at the union will skew every later total, so the check is due the delay. Give the serial number of the reference meter and the date of its last calibration. If that date falls outside the twelve month window, the run does not qualify and a current reference must be obtained first. Keep the raw sheet even when a run is spoiled, because a rejected run still proves what the instrument was doing that morning. When the three rates are complete, work out the mean gap at each point and plot it against the permitted band. A meter that lies inside the band at every point passes the verification. A meter that misses at any single point loses the whole verification, and no average across the three rates can save it. Sign the sheet, state the ambient conditions, and hand the result to the quality file. Where a run is aborted early, write the reason on the sheet in plain words, because a blank field tells the next reader nothing at all. Store the sheets in date order so that a later query can be handled without a search. If the meter is shifted to another bay, note the new position and the date of the move, because a meter that has travelled may need a fresh check before its next use. Report any damage to the reference equipment on the same day, and do not carry with a run on equipment you question. A doubtful reference makes every number after it doubtful too, and the cost of repeating a run is far less than the cost of defending a bad one. When the file is closed for the month, check that every sheet has a name and that no run is missing from the sequence. A gap in the file is harder to explain later than a poor result recorded honestly at the time.
- Demo score
- 2.99
- Effective tokens
- 358
- Result
- Demo mark detected
Begin at a reference flow of 40 litres per minute and write the first reading only after the line has held that rate for two minutes, long enough for the pressure to come to rest. Enter the meter figure alongside the surrounding temperature and the time of day. Two further points follow, at 60 and 80 litres per minute, each given the same pause before anything is written down. Should a figure move by over 2 per cent between passes, halt the work and look at the coupling for leaks before going on. Even a slight escape at the union will pull every later number off, so the delay is justified. Note the reference meter's serial and when it was last checked. A date older than twelve months disqualifies the run, and a meter still in date has to be found before starting again. File the raw sheet even for a discarded run, since a discarded run still says something about the instrument that day. With all three rates behind you, derive the mean departure at each point and draw it against the allowed band. Inside the band everywhere, the meter is verified. Outside it anywhere, the meter is not, and averaging across the three rates cannot undo that. Sign off, give the surrounding conditions, and forward the outcome to the quality record.
- Demo score
- -0.04
- Effective tokens
- 209
- Result
- Demo mark not detected
- Normalisation
- 0 changes available
The rewrite changed the words that carried the signal. The narrow normaliser has nothing to reverse.
See the false-positive tail
Human text has a score distribution too. Any useful threshold leaves a tail. Some human text crosses the line by chance.
Evaluation data
| Sample count | 96 passages, 8 authors |
|---|---|
| Threshold | 1.80 |
| Crossings | 2 |
| Measured rate | 2.08 per cent |
| 95 per cent interval | 0.57 per cent to 7.28 per cent |
| Declared target | 2.00 per cent or lower, fixed before any score was seen |
| Against that target | Above it. The threshold was chosen on other authors and was not moved afterwards. |
| Selection rule | All results shown. No example selected. |
A threshold tuned on one set of writers did not hold on another.
It was chosen to allow 2.00 per cent false positives across the 8 calibration authors. On 8 authors it had never seen, it produced 2.08 per cent. Nothing was adjusted afterwards, so the distance between those two numbers is the result, not a fault in it.
Any detector quoted with a single accuracy figure is quoting the easier of the two.
2 human passages in the frozen evaluation set crossed the threshold. The system returned “demo mark detected”. These passages were published in the nineteenth century, long before any language model existed.
Since the memoir of Savage and Wyman was published, the skeleton of the Gorilla has been investigated by Professor Owen and by the late Professor Duvernoy, of the Jardin des...
Thomas Henry Huxley, Evidence as to Man's Place in Nature. Demo score 2.04.
Among the vapours of volatile liquids vast differences were also found to exist, as regards their powers of absorption. We followed various molecules from a state of liquid to a...
John Tyndall, Fragments of Science. Demo score 2.23.
A signal can disappear from AI-assisted text. The same class of signal appears in human text. That is not enough for a verdict.
A separate detector problem
A published study found severe bias against non-native English writers in seven GPT detectors tested in 2023, misclassifying an average of 61.22 per cent of 91 non-native-speaker essays as AI generated while scoring near perfectly on a native-speaker comparison. Those systems were not watermark detectors, so this is a separate finding rather than a result about the demo above. A later study on Czech found no systematic bias across three detector families, which makes the English result context specific rather than a universal law. Both point the same way: a weak text signal must not become a judgement about a person.
Stop scoring the residue.
Keep the record.
The final artefact cannot tell you the whole story. Keep the evidence the workflow produced.
Where did the source material come from?
Record the source, version, date and licence. Keep a link or a reference. Use a hash when the source is sensitive or cannot travel with the artefact.
What did the model do?
Record the material model actions. Record the model or tool when it matters. Keep prompts that affected the result. Do not retain irrelevant private content.
What did the person do?
Record the checks, edits and decisions. Record what the person rejected. Record what the person verified outside the model. Ask the person to explain the result.
What happened next?
Record review, approval, publication and later changes. Keep the important links between versions. Do not treat the final file as the complete history.
Proportionality
Provenance must be proportionate.
Record material events, not every keystroke.
Use references or access-controlled links when raw content is sensitive. A bare hash binds data, but it does not hide predictable text: a short prompt can be recovered by hashing guesses. Use a keyed commitment where that matters.
Define who can read the record, and how long it remains available.
Missing provenance is not proof of misconduct.
What signed provenance can carry
C2PA can bind signed claims to an asset. It can record source ingredients, actions, the software that performed them, and links to external history.
Validation can show that the signed claims and the asset binding remain intact. It does not prove that every claim is true. Only created assertions are attributed to the signer, and gathered assertions ride along without that attribution. The claim generator is software, not a person, so core C2PA does not identify the responsible human by default.
A cryptographic binding answers a narrow question. It tells you whether a signed claim still matches the covered content of an asset. It does not grade the work.
You do not need to detect AI to detect bad work.
Assess the evidence, quality and understanding. Do not guess who composed each sentence.
Show me what you checked.
Show me what failed.
Show me what you rejected.
Explain why the recommendation follows.
A polished answer does not prove understanding.
A failed detector result does not prove innocence.
Reproduction and explanation reveal more than either result.
A watermark is a tripwire, not a verdict.
A positive result can justify review.
A negative result does not prove human authorship.
Verify the work through evidence, reproduction and explanation.