The numbers are in the announcement, the exam is not
The Dalian Institute of Chemical Physics, part of the Chinese Academy of Sciences, has announced version 3.0 Pro of the chemical-industry language model it develops with iFlytek and Alibaba Cloud. In the measurements the institute reports, text-based question-and-answer accuracy is 81.96 per cent and multimodal question-and-answer accuracy is 80.75 per cent, and those results are presented as improvements of 20.2 per cent and 31.4 per cent over version 3.0. The chemical-domain evaluation system the measurements rest on is named only in general terms.[1]
What an accuracy rate means depends on what the question set is. Without knowing which subfields the questions come from, how many there are, who wrote them and what counts as correct, neither the small gap between the two rates nor the reported jump over version 3.0 can be tested. The improvement percentages are also given on an overall score, and the announcement does not show that the overall score sits on the same scale as the two accuracy rates.[1]
Is execution a result or a design?
The announcement's main claim is that the model moves from answering knowledge questions to carrying out tasks. The structure described has four tiers: the large model comprehends professional knowledge, parses multimodal information and plans tasks; agents decompose the process, call tools, verify results and make adjustments around the task objective; below them sit professional skills, tools and application scenarios. That is the description of an architecture, and on its own it carries no measurement result.[1]
Question-and-answer accuracy cannot be the measure of that shift. Execution is measured by how many tasks finish end to end, how often a tool is called wrongly, at which step an error is caught, and which errors the verification step misses. In place of those measures the announcement cites more than 300 registered organisations and API calls past the 14 million mark; these are usage quantities and they do not show whether a task was carried out correctly. A weaker reading is also available: the institute may have run execution measurements and left them out because they do not fit an institutional announcement.[1]
What would test this number?
The institute says it will develop ChemELLM 4.0 with enhanced reasoning and multi-tool collaborative capabilities. If that announcement carries the evaluation set's name, its question count and its scoring rule together, the reported accuracy rates will become independently testable; if those details are withheld again, the numbers will remain results reported by the developer.[1]
Leaving the name out is not surprising. This text is an institutional announcement, and such texts often leave measurement detail to separate technical documents. Until that document appears, 81.96 per cent and 80.75 per cent stand as results reported by the developer rather than as a measure of capability in chemistry, and they cannot be compared against an independent test.[1]