Eigen RadarAI
Analysis

AI in expert work: Riemann progress and AMIE's video trial put validation first

Anthropic's unreleased model raised a known bound around the Riemann hypothesis while Google's AMIE was compared with physicians in simulated video consultations, putting human and formal validation beside the headline results.

Artificial Intelligence··Night
In a simulated clinical evaluation suite, a clinician seen from behind consults a patient actor behind glass by video while a second evaluator watches at unmarked controls.

One model, 650 paths, and an unfinished proof

TechCrunch's account of Anthropic's announcement says an unreleased model clearly raised the lower bound for which the Riemann hypothesis is known to hold. The result is not a general proof of the problem that has stood for more than 150 years, and the 1 million dollars prize remains unclaimed. The route to the result is at least as concrete as the result itself. An Anthropic employee without significant mathematical training asked the model to make a serious attempt, then let it work for a day and a half. The system tested 650 different ideas, distributed the work across 60 subagents, and produced 31 million output tokens. Two of Anthropic's in-house mathematicians confirmed the progress. The proof was then formalized with the open-source Lean proof assistant. That sequence joins broad machine exploration, expert review, and a formal check of the logical steps in one effort. The announced advance remains bounded by those checks; the model did not solve the Riemann hypothesis.[1]

How AMIE was tested on video

Google Research and Google DeepMind moved their research medical AI system AMIE into real-time video consultations. Built on Gemini and Project Astra with a multi-agent architecture, it is designed to interpret visual and auditory cues, guide a virtual physical examination, and reason diagnostically during the conversation. The evaluation did not take place in routine clinical care. It was a randomized study using simulated consultations with patient actors, with primary care physicians serving as the comparison group. Clinical evaluators rated AMIE favorably for the thoroughness of its history taking, diagnostic accuracy, the appropriateness of its management, and communication quality. The patient actors also preferred the video experience to text chat. Google's own account stops short of presenting those findings as a product ready for deployment. The company says AMIE remains a research system and that more work is needed before responsible real-world clinical use. The favorable video results therefore take their meaning from a defined simulation and comparison, not from broad use with patients in ordinary care.[2]

From the result to the validation chain

The two efforts operate in different expert domains and do not share a performance measure. The Riemann work aimed to advance a bound around an open mathematical problem, with Anthropic mathematicians and Lean providing the checks. The AMIE evaluation examined behavior in simulated video consultations against primary care physicians, with clinical evaluators and patient actors judging different parts of the experience. What connects them is that the model output was not left as the final word. Human review is decisive in both cases, though it performs a different job: checking a mathematical result in one and assessing clinical and communication behavior in the other. The formal proof assistant belongs only to the mathematical example. That distinction makes the validation chain part of the news whenever AI enters expert work. What the model did matters alongside the setting in which it acted, who compared the result, and which checks bounded the claim. Both systems are at noteworthy research stages; neither account announces broad, independently tested use in ordinary real-world practice.[1], [2]

References

  1. News sourceTechCrunchAn unreleased Anthropic model made progress on the Riemann hypothesis↩1↩2
  2. News sourceGoogleGoogle's AMIE system tried real-time video consultations with patients↩1↩2