Hcompany has launched Holo4, two AI models capable of operating software on computers and mobile devices through screen interactions, text entry, and code generation S¹. The 27B dense variant scores 61.7% on the OSWorld 2.0 computer-use benchmark, against 81.8% for Opus 5.5 S¹ (vendor's own figure; no error bars or released code provided). Hcompany's own chart excludes Holo4 from the efficiency frontier because the comparisons use different harnesses and task subsets across releases S¹.

My read: This is the first computer-use agent launch I've seen where the smaller dense model dramatically outperforms the larger MoE variant on the headline benchmark. I don't buy the "orders of magnitude fewer parameters" claim yet, because Hcompany doesn't quantify it against named competitors, and the OSWorld chart explicitly keeps Holo4 off the cost-performance line. The open-sourced trajectories are a genuine move, though the dataset viewer on Hugging Face is currently broken P². What I'd watch: whether Holo4's AutomationBench private-set score, which Hcompany says it will report once evaluated S¹, holds up against the public-set numbers it is publishing now.

Holo4 comes in two sizes: a 27B dense model and a 35B-A3B Mixture of Experts model S¹. Both are available on the H Models API S¹.

Hcompany additionally shipped Holotron4 Nano, a refreshed version of the Holotron 3 simulation environment S¹.

Holo4 can control software using graphical interfaces, code execution, the Model Context Protocol, and APIs S¹. The system operates across desktop platforms, web browsers, Android, sandboxed environments, and enterprise APIs using a unified model interface S¹.

Hcompany trained Holo4 through supervised and reinforcement learning on a large set of environments and tasks, including those generated by its Agentic Task Factory S¹. The models improve significantly over their Qwen base, according to Hcompany S¹ (vendor's own figure). The company tested Holo4 27B alongside Qwen3.8 27B on a FreeCAD 3D modeling assignment with identical prompts and testing setups S¹, but the final results are absent from the source material.

Benchmark scores reveal a 20-point gap between Holo4 variants

On OSWorld 2.0, Holo4 27B scores 61.7% and Holo4 35B-A3B scores 30.9%, compared with 81.8% for Opus 5.5 S¹ (vendor's own figure; no error bars or released code provided). The 20-point gap between the two Holo4 variants is unusual: the MoE model, which activates only 3B of its 35B parameters per token, trails the smaller dense model by more than half. Hcompany does not explain the gap in its release.

OSWorld 2.0 scores: Holo4 variants vs Opus 5.5

The OSWorld 2.0 cost-performance chart carries a caveat: models were tested under different harnesses and task subsets, and Holo4 is excluded from the non-dominated line connecting closed models S¹. On AutomationBench, Holo4 and its Qwen base models were evaluated using version 1.0.6 in Hcompany's internal harness S¹. Scores for competing models come from public-set scores and the official private-set leaderboard S¹. The comparisons therefore mix public and private evaluations. Hcompany says it will report Holo4's private-set score once it has been evaluated S¹.

Hcompany open-sources every trajectory behind its scores on public benchmarks S¹, though the Hugging Face dataset viewer for those trajectories is currently returning errors P². The previous generation, Holo3, was a vision-language model family optimised for GUI agents across web, desktop, and mobile P⁴, and Holo3.1 expanded that to mobile environments with native function-calling support P³.

For a team building automation workflows, the practical question is whether a 27B model that can drive a FreeCAD session or operate Android is worth deploying over a frontier API. The open-sourced trajectories let you audit the model's behaviour before committing. A QA engineer evaluating Holo4 for regression testing across a web app could replay those trajectories to check whether the model handles their specific UI patterns before paying for API calls. For a DevOps engineer, the availability of a 27B dense model that runs across desktop and Android means they can prototype automation scripts locally before scaling to API deployments.

Hcompany's next checkpoint is the AutomationBench private-set evaluation, which it says it will report once complete S¹.


Sources: S1 — Holo4: powering generalist computer-use agents · P2 — Hcompany/trajectories · Datasets at Hugging Face · P3 — Hcompany/Holo-3.1-4B · Hugging Face · P4 — Hcompany/Holo3-35B-A3B · Hugging Face · P5 — Qwen/Qwen3-30B-A3B · Hugging Face


Written from 5 sourced items, 5 of them primary.

More from Not A Tech Guy