What the detector can and cannot say

The Pew Research Center pushed nearly half a million English-language pages from the Common Crawl archive through a detector called Open Pangram. In a July 2026 sample, about 10 per cent of pages showed clear signs of AI authorship; limit the sample to pages published after ChatGPT's release and the share passes a third. The split by domain is sharper than the headline: roughly one .com page in ten carries the marker, against 4.6 per cent of .org pages and about 1 per cent of .edu and .gov pages.[1]

The limit sits in the tool. The write-up says plainly that detectors of this class make only a rough call on whether a person or a machine wrote a page, cannot say how much AI was involved or at which stage, and still misfire. That matters for a reader, because the interesting cases sit in the middle: a human draft tidied by a model, a paragraph filled in, an outline generated and then rewritten by hand. A single label flattens all of that into one bucket, and the 10 per cent inherits every judgement the detector made along the way.[1]

The number only Microsoft can read

The other approach in this week's news gives up on guessing. Xusheng Li, a developer at Vector 35, found that Microsoft Paint and Photos embed a 16-byte globally unique identifier into the pixels of images generated on the user's own machine. The prompt travels to Microsoft for moderation, and the identifier that comes back is encoded into the picture. Paint also sends the previous promptGenerationId as lastPromptGenerationId with its next request, so one prompt can be tied to the one before it.[2]

Seen from where the user stands, that precision points the wrong way. Li's argument is that Microsoft disclosed the watermark but did not make clear that the C2PA manifest carries an identifier linked to the user's prompts, and he compares it to the tracking dots colour laser printers have long added to pages. Microsoft did not respond to a request for comment. Whoever opens the image sees nothing, the person who made it sees nothing, and the party able to resolve the number back to an account is the party that issued it.[2]

The signals worth waiting for

Put the two side by side and the shared constraint is legibility. One method measures at web scale and admits its own error rate; the other writes an exact value and keeps it where only its issuer can read it. Neither hands the person reading a page or opening a picture anything they can check. The plausible alternative reading is that this is simply early: C2PA viewers could surface provenance in an ordinary file manager, and detector accuracy could improve enough to report a range instead of a verdict. Both are possible, and neither has arrived in the products people already use.[1], [2]

So here is the question worth holding: which of the two becomes checkable first? Two things would settle it. If Microsoft publishes what the identifier resolves to and how long it is retained, the watermark turns into a disclosure a user can weigh. If Pew or a comparable group publishes a version of this analysis with a stated error band and a split between fully generated and partly assisted text, the 10 per cent becomes a figure worth quoting. Until one of those lands, both numbers stay descriptions of the web in general, and what they say about the page in front of you stays limited.[1], [2]