The numbers the announcement carries
K-EXAONE 2.0 was published on Hugging Face under the Apache 2.0 licence as a model with 750 billion total and 37 billion active parameters, against 236 billion in the previous version. LG AI Research said it completed the architecture, data training, distributed computing and inference infrastructure on its own. In the measurements the company shared, the model averages 70.1 across 24 benchmarks spread over nine categories, presented as a 10% improvement on the previous version. On the OpenAI-MRCR long-context benchmark it scores 94.4, set against 71.5 for GLM-5.1. An average of 94.6 was reported on benchmarks assessing Korean social norms and ethical standards.[1]
Those two numbers do different work. An average across 24 benchmarks spread over nine categories compresses tests that measure different constructs into a single figure. The MRCR comparison names both the benchmark and the model compared against. The second is more informative because it is clear what is being set against what; but a single long-context retrieval test says nothing about what happens on the other 23.[1]
What the average does not carry
An average does not carry a distribution. The same value of 70.1 can come from a model clustered between 68 and 72 on every benchmark, or from one scoring 95 on some and 45 on others. Those are not the same system, and for a user the consequences are not the same either. The number of categories widens the problem: if nine categories were built to measure nine different constructs, the arithmetic mean across them corresponds to no quantity.[1]
The 10% improvement is likewise a comparison within its own family. Progress against a model's own previous version shows the development process worked; it does not show where the model stands against an outside baseline. It is also possible the average was chosen to convey a programme summary at home; reporting the output of a state-backed effort as a single figure is an ordinary communications choice, and in that case the average should be read as a summary rather than as the evidentiary claim.[1]
The methodological value of Apache 2.0
The most solid methodological element in the announcement is the licence. Because the weights are downloadable under Apache 2.0 and open to commercial use, the same 24 benchmarks can be re-run by an independent party and reported benchmark by benchmark. With a closed model that is not possible; there the reader is left with the producer's summary. Here the gap between summary and measurement can be closed.[1]
What is worth waiting for is therefore not another announcement but a distribution. If, by 30 November 2026, an independent evaluation publishes the same benchmark set benchmark by benchmark, the spread of the results will show how much the average represents: a narrow spread makes the mean a fair summary, a wide one means the announcement was being carried by the average. The same holds for the 94.6 on Korean norm benchmarks; what a test measuring a behavioural claim actually measures can be assessed only when the test and its scoring method are in the open.[1]