Paris-based H Company has released Holo4, a family of open-weight models built around a single wager: that real work crosses interfaces, and an agent that can click through a screen, run code in a sandbox and call an MCP server or a business API in one loop will beat one that can only do any single one of them.
The release covers two models. Holo4-27B is a dense vision-language model built on Alibaba's Qwen3.8 architecture; Holo4-35B-A3B is a mixture-of-experts variant on a Qwen3.6 base. H Company pairs both with its hai-agents harness, which feeds screenshots and tool results to the model, executes the clicks, keystrokes, code or tool calls it requests, and returns the results for the next turn — a reminder that in computer use the score belongs to the model-and-harness pair, not to weights alone. The company also updated Holotron4 Nano, a smaller sibling.
On OSWorld 2.0, a benchmark for agents driving real desktop software, H Company reports 61.7% for the 27B model and 30.9% for the larger MoE version. Its own table places the 27B figure next to a public 81.8% for Anthropic's Claude Opus 5.5 — a number H Company did not produce itself, and one drawn from a different harness and task set. The gap between the company's own two models, with the dense 27B beating the 35B-A3B by more than 30 points, is not explained in the announcement and is the first thing an outside evaluator will want to understand. Cost comparisons in the release likewise mix H Models API rates, Alibaba Cloud list prices and published pricing from OpenAI and Anthropic.
What H Company did do unusually well is make its claims checkable. It published 7,366 trajectories behind its public benchmark scores, replayable through its own viewer or downloadable from Hugging Face, and released the weights in FP16, FP8 and GGUF formats. Reviewers can inspect the sequence of actions rather than accept a number.
The licences are where the release gets more complicated than the marketing. Holo4-35B-A3B is Apache 2.0, but the better-performing Holo4-27B is released under CC BY-NC 4.0, which bars commercial use. Businesses that want the stronger model at self-hosted prices would need a separate arrangement — a split that puts performance and commercial rights in different columns of the same table.
That tension is the state of the category. Open-weight computer-use agents at this performance level remain rare, and a 27B model scoring in the low 60s on OSWorld 2.0 while running on hardware a company controls is a real option for teams that cannot send screenshots of their internal systems to a hosted API. Nobody has yet shown, though, that benchmark scores earned in curated environments — H Company's own examples include building a 3D model in FreeCAD and a game in Godot — survive contact with messy enterprise workflows. The published trajectories are a start. An independent rerun on a private task set is the test that would settle it.
Comments (0)
Log in to join the discussion
Log InNo comments yet