Eigen RadarAI
Analysis

Two validation boundaries outside the model's working view

Copilot processing an invisible instruction and coding agents failing to judge scientific correctness show why a model's working view needs separate user and domain validation.

Artificial Intelligence··Midday
Dark vault where a white shell reveals its violet filament inside a clear ring and connects to three externally tuned resonance bodies

The document surface diverged from Copilot's input

The attack demonstrated by security researcher Hakon Maloy against Copilot inside Word depends on the model and reader receiving different representations of the same document. Instructions are placed in a tiny font and coloured white on white, leaving the text invisible to the user. Copilot removes colour and font-size information before processing, making the words visible again in the model input. The system executes the instruction, copies it into a newly created document, and can extend the chain when that output becomes input to a later Copilot operation. Financial figures changed without the user noticing in the demonstration. The issue was reported to Microsoft's Security Response Center on 6 March, confirmed on 31 March, and followed by two mitigation steps. The researcher disclosed the work on 28 July after reproducing the attack against GPT-5.6, while withholding the wording of the payload. The report gives no count of real users and describes no field deployment. Its central boundary is that the document surface visible to a user does not by itself describe the content the model actually processes.[1]

Working code did not establish its own scientific acceptance test

A field report by OpenAI and academic partners records another representation gap across eight scientific-software cases. Agents delivered substantial runtime gains in modernisation and optimisation work: RustQC fell from 15 hours 34 minutes to 14 minutes 54 seconds, while HelixForge ran 59.6 times faster on GPU tasks and 98.6 times faster in its main compute step. Runtime fell by about 31 percent for HI.SIM and about 15 percent for hifiasm. Some projects also measured output agreement; rustar-aligner matched the original STAR output for 99.815 percent of single-end reads and 99.883 percent of paired-end reads. Even so, the report says the agents could not reliably assess whether their own work was scientifically correct. Inverted parameters and calculation errors in the bayesm case appeared during rigorous statistical calibration tests. People established the validation criteria and definition of success, with the test changing according to each project's domain context. The authors also state explicitly that the eight cases are not a representative evaluation.[2]

Transformation and acceptance occur at different checkpoints

The reports concern different risk domains: the Copilot case examines input visibility in a controlled security demonstration, while the scientific-software report examines domain correctness in code output. Their shared constraint is that the representation on which a model operates cannot establish the final acceptance boundary on its own. When Word strips colour and font size, content invisible to the user enters the processing layer. In scientific code, an output that compiles or runs faster does not automatically define the statistical and scientific behaviour a project requires. Validation is therefore established through two external relationships. The first must inspect the difference between the surface a user sees and the model input. In the second, domain specialists set the reference output, calibration test, and threshold for success. Treating the findings as one security category would be inaccurate: propagation of a hidden instruction and detection of an inverted parameter are different mechanisms. Together they offer a practical frame in which the human-visible view and domain test must be designed alongside the object transformed by the model.[1], [2]

References

  1. News sourceThe DecoderA Copilot attack that spreads through instructions hidden in Word documents was demonstrated↩1↩2
  2. News sourceThe DecoderOpenAI's field report says coding agents cannot check scientific correctness↩1↩2