> ## Content Index
> Fetch the complete content index at: https://www.notatechguy.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Holo4 27B scores 61.7% on OSWorld 2.0, trails Opus 5.5
- URL: https://www.notatechguy.com/holo4-27b-scores-61-7-on-osworld-2-0-trails-opus-5-5/
- Published: 2026-09-28T16:14:47.000Z
- Updated: 2026-09-28T16:14:48.000Z
- Description: Hcompany's Holo4 agentic models click, type and code across desktops and phones, but benchmark comparisons mix harnesses and task sets.
- Author: Marcello Babbili
- Tags: Technology & AI

Hcompany has launched Holo4, two AI models capable of operating software on computers and mobile devices through screen interactions, text entry, and code generation [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com). The 27B dense variant scores 61.7% on the OSWorld 2.0 computer-use benchmark, against 81.8% for Opus 5.5 [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com) (vendor's own figure; no error bars or released code provided). Hcompany's own chart excludes Holo4 from the efficiency frontier because the comparisons use different harnesses and task subsets across releases [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com).

**My read:** This is the first computer-use agent launch I've seen where the smaller dense model dramatically outperforms the larger MoE variant on the headline benchmark. I don't buy the "orders of magnitude fewer parameters" claim yet, because Hcompany doesn't quantify it against named competitors, and the OSWorld chart explicitly keeps Holo4 off the cost-performance line. The open-sourced trajectories are a genuine move, though the dataset viewer on Hugging Face is currently broken [P²](https://huggingface.co/datasets/Hcompany/trajectories?ref=notatechguy.com). What I'd watch: whether Holo4's AutomationBench private-set score, which Hcompany says it will report once evaluated [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com), holds up against the public-set numbers it is publishing now.

Holo4 comes in two sizes: a 27B dense model and a 35B-A3B Mixture of Experts model [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com). Both are available on the H Models API [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com).

Hcompany additionally shipped Holotron4 Nano, a refreshed version of the Holotron 3 simulation environment [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com).

Holo4 can control software using graphical interfaces, code execution, the Model Context Protocol, and APIs [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com). The system operates across desktop platforms, web browsers, Android, sandboxed environments, and enterprise APIs using a unified model interface [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com).

Hcompany trained Holo4 through supervised and reinforcement learning on a large set of environments and tasks, including those generated by its Agentic Task Factory [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com). The models improve significantly over their Qwen base, according to Hcompany [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com) (vendor's own figure). The company tested Holo4 27B alongside Qwen3.8 27B on a FreeCAD 3D modeling assignment with identical prompts and testing setups [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com), but the final results are absent from the source material.

Benchmark scores reveal a 20-point gap between Holo4 variants

On OSWorld 2.0, Holo4 27B scores 61.7% and Holo4 35B-A3B scores 30.9%, compared with 81.8% for Opus 5.5 [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com) (vendor's own figure; no error bars or released code provided). The 20-point gap between the two Holo4 variants is unusual: the MoE model, which activates only 3B of its 35B parameters per token, trails the smaller dense model by more than half. Hcompany does not explain the gap in its release.

![OSWorld 2.0 scores: Holo4 variants vs Opus 5.5](https://storage.ghost.io/c/6e/89/6e896869-22ef-4281-a213-b4c462c17cff/content/images/2026/09/chart_c730099fb7a46fba19d4.png)

The OSWorld 2.0 cost-performance chart carries a caveat: models were tested under different harnesses and task subsets, and Holo4 is excluded from the non-dominated line connecting closed models [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com). On AutomationBench, Holo4 and its Qwen base models were evaluated using version 1.0.6 in Hcompany's internal harness [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com). Scores for competing models come from public-set scores and the official private-set leaderboard [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com). The comparisons therefore mix public and private evaluations. Hcompany says it will report Holo4's private-set score once it has been evaluated [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com).

Hcompany open-sources every trajectory behind its scores on public benchmarks [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com), though the Hugging Face dataset viewer for those trajectories is currently returning errors [P²](https://huggingface.co/datasets/Hcompany/trajectories?ref=notatechguy.com). The previous generation, Holo3, was a vision-language model family optimised for GUI agents across web, desktop, and mobile [P⁴](https://huggingface.co/Hcompany/Holo3-35B-A3B?ref=notatechguy.com), and Holo3.1 expanded that to mobile environments with native function-calling support [P³](https://huggingface.co/Hcompany/Holo-3.1-4B?ref=notatechguy.com).

For a team building automation workflows, the practical question is whether a 27B model that can drive a FreeCAD session or operate Android is worth deploying over a frontier API. The open-sourced trajectories let you audit the model's behaviour before committing. A QA engineer evaluating Holo4 for regression testing across a web app could replay those trajectories to check whether the model handles their specific UI patterns before paying for API calls. For a DevOps engineer, the availability of a 27B dense model that runs across desktop and Android means they can prototype automation scripts locally before scaling to API deployments.

Hcompany's next checkpoint is the AutomationBench private-set evaluation, which it says it will report once complete [S¹](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com).

---

*Sources: [S1 — Holo4: powering generalist computer-use agents](https://huggingface.co/blog/Hcompany/holo4?ref=notatechguy.com) · [P2 — Hcompany/trajectories · Datasets at Hugging Face](https://huggingface.co/datasets/Hcompany/trajectories?ref=notatechguy.com) · [P3 — Hcompany/Holo-3.1-4B · Hugging Face](https://huggingface.co/Hcompany/Holo-3.1-4B?ref=notatechguy.com) · [P4 — Hcompany/Holo3-35B-A3B · Hugging Face](https://huggingface.co/Hcompany/Holo3-35B-A3B?ref=notatechguy.com) · [P5 — Qwen/Qwen3-30B-A3B · Hugging Face](https://huggingface.co/Qwen/Qwen3-30B-A3B?ref=notatechguy.com)*

---

*Written from 5 sourced items, 5 of them primary.*

## More from Not A Tech Guy

- [LLVM trends on GitHub as AI agents target compiler code](https://www.notatechguy.com/llvm-trends-on-github-as-ai-agents-target-compiler-code/)
- [AI security tool reverse-skill hits 38,000 GitHub stars](https://www.notatechguy.com/ai-security-tool-reverse-skill-hits-38-000-github-stars/)
- [GRPO step-level bias: GRAFT graph method claims gains on agent tasks](https://www.notatechguy.com/grpo-step-level-bias-graft-graph-method-claims-gains-on-agent-tasks/)