Eigen RadarAI
Analysis

Holo4 aims to work across screens, code and APIs with one model

H Company released two Holo4 models designed to act through screens, code, MCP tools and APIs, and posted their weights. The central claim is that one model can continue a task across interfaces. The company’s benchmark comparisons use different testing conditions. Published action traces let others inspect the disclosed runs, but they do not establish reliability across other workflows.

Artificial Intelligence··Evening
A developer works across a laptop, tablet and second monitor.

One model reaches four ways to act

H Company released two Holo4 AI agent models for acting in software. TPS also reported the September 28 release and the design shared by both models: they can work through graphical interfaces, code, Model Context Protocol (MCP) tools and application programming interfaces (APIs). In the company’s design, a task can move from a screen control to code execution or a service call without selecting a separate specialist model. The dense version has 27 billion parameters; the mixture-of-experts version has 35 billion in total. Both are offered through the H Models API, and their weights are posted on Hugging Face. Those forms of access do not measure whether tasks finish reliably in real settings.[1], [2]

Training tasks span multiple interfaces

H Company says it combined supervised learning with reinforcement learning. Its internal task factory generated about 10,000 assignments involving web applications, desktop programs and MCP servers. In some environments, the same state can be reached through a screen or an MCP tool, forcing the model to choose a route. The company also changed its execution harness to include memory intended to track state over hundreds of steps and shell access on the desktop machine. Preserving task context as an agent switches interfaces is part of the use case the design targets. H Company says it has posted downloadable and replayable action traces for public benchmark tasks. Those traces permit inspection of the disclosed runs rather than merely their aggregate score.[1]

Public traces do not settle real task performance

In H Company’s OSWorld 2.0 comparison, the dense Holo4 version scored 61.7 percent and the mixture-of-experts version 30.9 percent. TPS repeats those figures from the company’s release; it does not offer a separate comparison test. Figures for other models in the announcement come from different sources and conditions, so they are not results from one matched harness. The AutomationBench measurements used the company’s internal setup, and it has not reported a result on the private task set. Published action traces can show which steps the models took on specified public tasks. The announcement does not establish how reliably the same model handles changing screens, permissions and interruptions in an organization’s own software.[1], [2]

References

  1. News sourceH CompanyH Company releases Holo4 agents that work across software interfaces↩1↩2↩3
  2. News sourceTPSHolo4 models launch with access to multiple computer interfaces↩1↩2