Eigen RadarAI
Analysis

Anthropic adds watermarks as researchers rebuild prompts from outputs

Anthropic began watermarking Claude text for European Union rules, while a preprint reports reconstructing short prompts from model outputs. Together, they sharpen the question of what provenance labels can actually establish.

Artificial Intelligence··Evening
A translucent sheet covered in fine rings sits beside three irregular fragments whose colored traces pass through a glass prism onto a pale card.

An invisible watermark also marks edited text

Anthropic said in a support article that it has begun embedding machine-readable watermarks in text Claude produces, in order to satisfy the European Union’s AI Act. The mark does not stop at text the model produced; it also goes on content the model merely edited. The law covers models released after 2 August and gives providers until December 2026 to update earlier ones. The watermark in text output is invisible to the user, while other generated files receive digitally signed provenance metadata where supported and non-text content gets C2PA metadata. Anthropic says every new model it offers globally marks generated content from day one. By the company’s own wording a detected mark “provides a signal that content was processed by Claude, but is not fully conclusive.” The detection tool the law requires has not yet been released, so how thoroughly the mark holds cannot be tested from outside. That design stretches the provenance label onto every edit the model touches, not only fresh generation. Readers still see the same sentence while a verification check looks for a machine-readable signal, and without the detection tool outsiders cannot measure recovery rates.[1]

An inverse model rebuilds short prompts from answers

Researchers at IIT Bombay and Adobe Research describe an inverse language model that predicts the previous token instead of the next one, and reconstruct the prompt behind a model’s output. In the tested examples the recovered prompts matched the originals word for word. The method, called Previous-Token Prediction, trains on synthetic data generated from the target model using only its output text. An inverse model built on the smaller Qwen-3-0.6B recovered prompts from GPT-4o output. The preprint, posted as arXiv 2607.29378, has not been peer reviewed, and the authors tested only prompts of one or two sentences, leaving long multi-paragraph system prompts untried. The work claims that output text can leak a hidden instruction, so the answer a reader sees may carry a short user prompt almost intact. The method does not need the target model’s weights; it builds synthetic training data from text output alone. That narrow test set narrows the claim of success, and without peer review the results remain preliminary. Even so, word-for-word matches on short prompts show that an output can carry a trace of the input.[2]

What can a provenance label prove?

Watermarking and prompt reconstruction make two different limits of provenance labels visible in the same week. Anthropic’s mark supplies a signal that Claude processed the text, yet by the company’s own wording it is not fully conclusive; because the mark also lands on edited prose, it blurs pure generation and mere touch-up. Without the detection tool, outsiders cannot verify recovery rates either. At the other end, word-for-word recovery of short prompts from model answers shows that an output can carry a trace of the input. The two tracks are not substitutes: a watermark is a provider-applied mark, while prompt reconstruction is an estimate extracted from the output. Together they still force the question of what a label actually proves. Saying Claude processed a passage does not reveal the prompt; recovering a prompt does not show whether legal watermarking is present. European Union rules push providers to mark content, while the research reminds readers that the output itself may leak a hidden input. Provenance therefore becomes a question of scope, certainty and the input trace an answer may still carry.[1], [2]

References

  1. News sourceArs TechnicaAnthropic's new watermark also marks text Claude only edited↩1↩2
  2. News sourceThe DecoderA model running backwards rebuilds the prompt from the answer↩1↩2