What was measured

Echoverse, published by Microsoft Research, is built on a small number of deep, stateful environments that change over time rather than many shallow tasks: ten domain worlds and two capability worlds. Code, datasets and graded task verifiers for four complete environments were shared on GitHub and Hugging Face. On the published measurements Qwen3.5-9B rises from 36.5% to 67.1% across all environments, while GPT-5.4 sits at 80.7% as the reference model in the same comparison. The verifier grading the tasks is GPT-4.1.[1]

Those two measurements answer different questions. The rise inside the training environments shows how well the model fits the distribution it was trained on. The rise outside them is what speaks to whether the method teaches a general competence. The value of a training system is settled by the second measurement: the first says the design works, the second says it is useful.[1]

The three numbers from outside

On environments not seen in training, reinforcement learning lifted performance from 58% to 69%: a gain of 11 points. On two fully external benchmarks the picture narrows: 66.5% to 71.5% on WebVoyager and 40.5% to 43.4% on Online-Mind2Web.[1]

A gain of 11 points on held-out environments against 5 and 2.9 on familiar benchmarks fits a system teaching interaction habits that travel within its own family of environments, a habit that works on tasks built from the same construction rather than a general competence at using a computer. The competing account is that the two external benchmarks carry different task definitions and approach saturation at their upper ranges, in which case the small gain marks the benchmark's limit rather than the method's.[1]

What makes replication possible

The experiment that would settle this sits inside the released material. Since the graded verifiers and datasets were shared, the same outputs can be re-scored with a verifier from a different family. That is how one tests whether GPT-4.1 favours behaviour from its own family, and how one separates how much of the 67.1% comes from the method and how much from the scorer. The technical report is not work that has passed conference review; the results should be read as the company's own measurements.[1]