Eigen RadarAI
Analysis

Reasoning tokens do not pay off on every task as Copilot makes effort adjustable

A preprint published on arXiv on 29 August proposes the Token Economy Score to measure returns on reasoning tokens; step-by-step tasks scored more efficiently than knowledge-recall sets. With GitHub's August update, Copilot in Visual Studio lets developers set thinking effort to low, medium or high. Research questions efficiency while the product exposes a direct effort dial.

Artificial Intelligence··Midday
Close-up of a steel balance on a bright jeweller's bench, with glowing fine coils on one pan and a gear on the other as a hand adjusts a slider beside three beam notches.

The Token Economy Score measures task shape

The authors propose the Token Economy Score, which divides a reasoning model's accuracy gain over a non-reasoning baseline by how many more tokens it generated. Across 151 model-benchmark evaluations on seven benchmarks in mathematics, code generation and science reasoning, the shape of the task predicted efficiency better than its difficulty: step-by-step problems such as AIME 2025 scored well, while knowledge-recall sets such as MMLU-Pro scored lower. The authors also report systematic diminishing returns as the reasoning effort setting is raised, and cases in which more thinking lowered accuracy. Their recommendation is to switch reasoning on by task type, effort level and deployment context rather than leave it on everywhere. The preprint went up on arXiv on 29 August and was accepted at the 2026 TPCTC conference, and the score is offered as a way to compare whether extra reasoning tokens buy accuracy on a given benchmark rather than treating longer chains as uniformly valuable.[1]

Copilot split thinking effort into three levels

With GitHub's August update, organisations can publish custom Copilot agents across repositories and have Visual Studio detect them automatically. A developer can set the model's thinking effort to low, medium or high, trading reasoning depth against token consumption. The update also brings usage details and plan information into the context window and warns as a limit approaches. Favourite models can be pinned and unused ones collapsed, with model capabilities, context window sizes and cost information in the same place. The Git agent reviews uncommitted changes and commits before a pull request is opened and shows its findings inline in the editor. The features are available on the Free, Student, Pro, Pro+, Max, Business and Enterprise plans, and organisation-wide custom agents can now be detected automatically in Visual Studio without per-repository manual registration. The reported findings come from the authors' own runs, and the work is presented as a preprint that has not been peer-reviewed.[2]

Measurement and product answer the same question at different layers

The preprint argues that extra reasoning tokens do not return equally across benchmark suites, so enterprise deployments should tune effort by task. The Copilot update brings that argument to product level: a developer can pick low, medium or high effort to control token spend directly, with usage detail visible in the context window and warnings as limits approach. Both developments carry the same economic tension — the cost of deeper reasoning — but one through academic metrics across 151 evaluations and the other through an in-IDE dial plus inline Git review before pull requests open. The authors saw diminishing returns as reasoning effort rose; Copilot exposes effort beside plan and cost information in the same context window. For readers the concrete news is a efficiency score publishing alongside a thinking-effort control arriving in a developer tool the same week. The reported findings come from the authors' own runs, and the work is presented as a preprint that has not been peer-reviewed.[1], [2]

References

  1. News sourcearXivThree researchers put a price on reasoning tokens and find where they stop paying↩1↩2
  2. News sourceGitHub ChangelogCopilot in Visual Studio gains organisation-wide custom agents and a thinking-effort dial↩1↩2