One model, two runs
In the Deep Research evaluation Thomson Reuters ran itself, the same model was measured twice. Working with web access alone, Thomson's factual accuracy stayed at 0.53 while GPT-5.4 scored 0.65 under the same condition. Once access to the company's own content was opened, Thomson rose to 0.83 and GPT-5.4 stayed at 0.82.[1]
For a builder, those two runs pull apart two things that a single score usually blends: the model's general capability and the archive it can reach. Andrew Bean, who led the evaluation, attributes the gain to being able to train and practise on the company's own tools. It is worth saying that this reading is not the only one available: the retrieval interface may also differ between the two runs, so part of the swing could come from how those in-house tools are built rather than from the archive itself.[1]
The ledger behind the build
The shape of the spending is instructive too. About $40 million went on staff and compute over more than 2 years, while the final training run itself came to $450,000. The base is Alibaba's open Qwen3.5-397B. An intermediate version called Snowdon was retrained with Imperial College for safety, ethics and political neutrality. Training rests on the Westlaw, Practical Law, Checkpoint and Reuters archives, with hundreds of full-time domain experts involved.[1]
Chief technology officer Joel Hron and research chief Jonathan Schwartz give three reasons for taking this route rather than fine-tuning a leading model: fine-tuning tends to erode general capability, it ties the company to a provider's inference costs, and building something you own compounds over time. For now the model runs in one place, the Tabular Analysis feature inside CoCounsel Legal, which is document review, the kind of work where a small cheap model makes economic sense. That fits the measure argued in this column on 15 August: the deciding number for a builder is cost per completed task at a stated effort level rather than the advertised token price.[1], [2]
What outsiders can check
On general benchmarks the picture is plainer. Thomson scores 0.823 on Stanford LegalBench, behind Gemini 3.1 Pro and GPT-5.5, and sits just behind Opus 4.8 on the Harvey Legal Agent benchmark. Bean describes the model as within the range of the others but not yet the leader. All of these numbers come from the company's own testing. The company also says it has used less than 10 percent of the content available to it.[1]
That makes the planned open release the real test. The company is considering putting a small version on Hugging Face under a non-commercial licence. If that version appears by 30 November 2026 and independent outside evaluations place it close to the reported LegalBench level without access to the archive, the reading that the archive carries most of the gain gets stronger. If those evaluations put the model on its own clearly higher, the weight belongs with the training method instead.[1]