Which moment does the evaluation measure?
TutorMoments replays real tutoring transcripts and stops at the moments where a tutor has to choose: ease the problem and move the student along, or push the student toward deeper reasoning. The dataset is 462 de-identified transcripts from one-to-one mathematics tutoring in the United States, covering grades two through seven. In total, 27 teachers marked more than 1,500 key moments; the annotations contain 738 scaffolding and 260 rigor moments.[1]
Seven models were tested with two prompts: one plain, the other spelling out the trade-off being measured. The scores human tutors received were published too: 0.458 for scaffolding, 0.182 for rigor and 0.496 for avoiding over-scaffolding. That third figure shows how hard the task is, because a human tutor does not score high on this scale either.[1]
Why does the prompt raise the score?
Every model scored higher under the prompt that names the trade-off than under the plain one. If telling a model what is being measured raises the measurement, part of the score becomes a property of the instruction rather than of tutoring ability. There is a serious answer to that: the explicit prompt may remove an ambiguity a real tutor never faces. In that case the plain-prompt score is the pessimistic bound. Separating the two means running intermediate conditions that vary only the level of detail in the instruction, not its content.[1]
What the work cannot establish is equally clear. The tech report has not been through peer review, and replaying a transcript is not a prospective trial with real students; no learning outcome is measured. The annotations rest on the judgement of 27 teachers, so the standard represents their consensus. Two steps would raise confidence: rescoring the same moments with an independent set of annotators, and a prospective study measuring whether a model's choice actually changes what a student does next. Because Ai2 released the dataset and code openly, the first step can be taken today.[1]