H Company's Holo4 Handles Every Interface a Computer Offers, in One Model

Holo4 can click menus, write code, and call APIs within a single task. For businesses weighing agent costs, that breadth at a smaller scale is the real story.

AI2Day NewsdeskEditor: Lee Brown3 min read
A desktop computer screen displaying multiple open software windows simultaneously, including a 3D modeling application, a code editor, and a web browser, photo
Share

Key points

  • Holo4's 27B model scores 61.7% on OSWorld 2.0, a benchmark for computer-control tasks, trailing only the largest closed models while using far fewer computing resources.
  • Unlike most AI agents, Holo4 switches between graphical interfaces, code execution, and software APIs within one task without switching models.
  • H Company is publishing every recorded run behind its benchmark scores so anyone can replay and verify them.
  • The release also includes Holotron4 Nano, built on Nvidia's Nemotron Nano Omni base, extending the same agent training to a smaller, cheaper model.

Most AI agents, the software programs that carry out multi-step tasks on a computer without constant human input, have a narrow skill set. A model trained to click through graphical menus goes blank when there's no screen to look at. One trained to call software APIs, the coded connections that let programs talk to each other, gets stuck in front of an app that has no such connection. Real office work doesn't respect those limits.

H Company's Holo4, released this week and first reported by Hugging Face, comes in two versions: a 27-billion-parameter model and a 35-billion-parameter model using a Mixture of Experts design, where only part of the model activates for any given task. Both are available through H Company's own API.

What does Holo4 actually do?

It handles all four major ways software exposes itself: graphical point-and-click interfaces, written code, MCP tools (a standard that lets AI models call outside services), and direct software APIs. One task can require all four, and Holo4 moves between them without being reconfigured.

H Company demonstrated this with two tasks. Building a detailed 3D model of the Eiffel Tower in engineering software took 84 agent calls and 1.3 million tokens of processing. A working Pac-Man game that plays itself took 68 calls and 2.4 million tokens. The same Pac-Man task given to Qwen3.8 27B, the base model Holo4 was trained on, needed 197 calls and 11.4 million tokens. Fewer steps and fewer tokens means lower cost per task.

How does it compare to the big players?

On OSWorld 2.0, Holo4 27B scores 61.7%. Anthropic's Opus 5.5 scores 81.8%. That gap matters. But Holo4 reaches its score at a fraction of the parameters and a lower price per run, which is the practical comparison for a business deciding what to deploy.

Model OSWorld 2.0 Score Type
Opus 5.5 81.8% Closed, large
GPT-5.6 Sol 66.2% Closed, large
Holo4 27B 61.7% Open, small
Holo4 35B-A3B 30.9% Open, MoE

Task subsets and evaluation setups differ across these results, so the numbers are directional, not a controlled race.

What H Company does with its scores matters as much as the scores themselves. Every recorded agent run behind its public benchmark numbers is published, replayable at trajectories.hcompany.ai or downloadable from Hugging Face. On 22 September we reported that the UK AI Security Institute took a similar step, publishing verified test results under an open standard so researchers could repeat the work. A lab releasing full run logs is still the exception.

Should you worry about whether these results hold up?

Training used roughly 10,000 tasks built by an internal pipeline H Company calls the Agentic Task Factory, which generates verifiable tasks from software documentation and screenshots. H Company also rebuilt its control loop, the software that feeds the model's actions back into the environment over hundreds of steps and tracks the task across that full span.

For workers whose employers are evaluating AI agents, the practical question is whether a model trained on broad computer-use tasks holds up on your specific software. Open run logs let technical staff verify the benchmark claims rather than take them on faith. That's worth more than any single score.

© 2026 AI2Day