🔍 Read the full analysis: Holo4: Powering Generalist Computer-use Agents on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
H Company has released Holo4, a pair of open-weight models designed to operate software through graphical interfaces, code and tool calls. The company reports a 61.7% OSWorld 2.0 score for the 27B model, but the results have not been independently verified and comparisons use different evaluation setups.
H Company has released Holo4, a series of open-weight agent models built to operate software through screens, code and tool interfaces, as described in the original analysis. The company says its 27-billion-parameter model scored 61.7% on OSWorld 2.0, a computer-use benchmark, putting it below the 81.8% score it attributes to Opus 5.5; outside evaluators have not yet confirmed the comparison.
Holo4 is available in two versions: a 27B dense model and a 35B-A3B Mixture of Experts model. H Company says both can click and type in graphical interfaces, write and run code, and call MCP or API tools, selecting an interface according to the task. The models are offered through the H Models API and as downloads on Hugging Face in FP16, FP8 and GGUF formats.
The company reports OSWorld 2.0 scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. It compares the 27B result with an 81.8% score for Opus 5.5, which it identifies as the strongest closed model in the comparison. The announcement does not explain why the larger Holo4 variant scored substantially lower than the dense model.
H Company says it trained the models with supervised and reinforcement learning across tasks and environments, including tasks generated by its Agentic Task Factory. It has also published trajectories behind its public benchmark results on its website and made them downloadable from Hugging Face. Those records could allow researchers to inspect how the models carried out benchmark tasks, though they do not by themselves validate the reported scores.
One Model Across Software Interfaces
Many workplace tasks combine actions across a screen, code and business tools. A system limited to one interface may be unable to continue when a task moves from clicking through an application to running a script or calling an API. H Company’s central claim is that Holo4 can work across those interfaces with the same model, which could simplify how developers build agents for mixed workflows.
If independent tests reproduce the reported results, open weights could give businesses another way to run computer-use agents, including through self-hosting. H Company also says Holo4 delivers performance at lower cost than frontier closed models. That cost comparison depends on the company’s stated pricing assumptions and evaluation setup, so readers should treat it as a vendor comparison rather than a settled measure of operating costs.
The public trajectories add a practical avenue for scrutiny: outside teams can examine individual task attempts rather than relying only on a score. The value of that release will depend on whether independent evaluators can reproduce the results under comparable conditions and establish how well performance carries over to ordinary business software.
From Holo Models to Holo4
Holo4 follows H Company’s earlier Holo agent model and arrives alongside an updated Holotron line, called Holotron4 Nano. The company’s benchmark notes identify Qwen3.8 27B as the base for the dense Holo4 and Qwen3.6 35B-A3B as the base for the MoE version. H Company says side-by-side examples in FreeCAD and Godot show improvements over the respective base models, using the same prompts and harness.
The benchmark comparisons have qualifications. H Company says its OSWorld results and several cost calculations use company-selected releases, harnesses and task subsets that may differ from those used for other models. For AutomationBench, Holo4 was tested in the company’s internal harness, version 1.0.6, while comparison scores for other models come from the public set. The associated cost figures draw on a leaderboard using the private set, so the score and cost comparisons do not share the same evaluation basis.
The company’s cost estimates also rely on specific pricing inputs: Holo4 API rates for one run, Alibaba Cloud list prices for Qwen, and launch data for OpenAI and Opus effort sweeps. H Company says the assumptions include a cache-hit price of 20% of input for the MoE model’s Qwen cost. These figures are estimates tied to those assumptions, not a universal cost for deploying each model.
“Real work is not siloed that way, and a single business task can require combining these different approaches.”
— H Company
Benchmark Questions Still Open
All headline scores cited here are company-reported; independent reproductions are not included in the announcement. Differences in model releases, harnesses and task subsets make direct comparisons difficult, and H Company itself flags those differences. Publishing trajectories offers material for review, but the available information does not establish that third parties have reproduced the results.
Holo4 has not yet been evaluated on the AutomationBench private set, according to the company. It says it will report that result when the evaluation is complete. The reason for the large gap between the two Holo4 variants on OSWorld 2.0 is also not given. The announcement provides demonstrations in professional software, but does not establish reliability across a broad range of real business workflows.
Private Benchmark and Replications
H Company says it plans to publish Holo4’s AutomationBench private-set result after that evaluation is finished. Independent benchmark submissions and reproductions could then test how the public OSWorld results hold up under other evaluators’ methods. Developers can access the models through the H Models API or download them from Hugging Face; whether they will perform consistently outside the reported tests remains an open question.
Key Questions
What is Holo4?
Holo4 is a series of open-weight agent models from H Company. The company says the models can operate software through graphical interfaces, code, MCP and APIs.
What Holo4 models are available?
H Company released a 27B dense model and a 35B-A3B Mixture of Experts model. Both are available through the H Models API and for download on Hugging Face in several formats.
How did Holo4 score on OSWorld 2.0?
H Company reports scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. These are company-reported results; independent confirmation is not provided in the announcement.
Can the benchmark results be independently checked?
H Company has released trajectories behind its public benchmark scores, which researchers can inspect. Independent reproduction of the reported scores remains pending, and the company says its comparisons involve differing releases, harnesses and task sets.
What evaluation is still pending?
H Company says Holo4 has not yet been evaluated on the AutomationBench private set and plans to report results after that test is complete.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
