The type of a written step separates in the middle layers
Researchers at KAIST and Naver AI Lab asked whether the reasoning steps a language model writes on the page can also be told apart in its hidden numerical representations. KAIST and Naver AI Lab defined eight recurring operation types, including extraction, decomposition, formula recall, deduction and computation. Qwen and Gemma models solved maths problems; the traces were split into segments, and GPT-5 labelled each segment with one of those operations.[1]
Across all three models the operation types could be told apart in the hidden states, with the clearest signal in the middle layers. A classifier that looked only at the tokens used did worse than one that read the internal states. Position along the solution path did not explain the effect either. The same function word acquired a different internal representation in middle and later layers depending on which operation surrounded it.[1]
The type of the step is not the same as a correct result
Blocking attention to the preceding 30 tokens weakened the operation signal for that step; the type does not form in isolation, it leans on prior context. Even on incorrectly solved problems, the kind of step the model was taking — computing, retrieving a formula, or deducing — stayed identifiable inside. A flawed calculation still looked like a calculation internally, even when the result was wrong.[1]
The separability also appeared in Llama-3-8B, and classifiers trained on Qwen3-8B transferred to GPQA-Diamond and MATH-500. The experiments are limited to maths tasks and a handful of models. Catching errors or steering a model mid-generation is not what this measurement shows; that claim needs a separate design.[1]
The written trace is not a test of correctness
This evidence makes it harder for me to treat written reasoning as empty display: the operation type leaves a mark in the middle layers that goes beyond surface wording. The same evidence also stops me from treating that mark as a test of correctness. A 'computation' step on the page does not show that the arithmetic held; it shows that the model was performing a computation-type operation at that moment.[1]
The next evidence to look for is whether an operation-type readout can separate a wrong result from a right one: a run, on the same maths traces, in which an independent hand tells failed calculation steps from successful ones. Without that split, the type of the written step is only a reading of what the model takes itself to be doing.[1]