Senior Software Engineer - Paris
Job Description & Overview
When we released Holo4 on 28 September, we published every trajectory behind its public benchmark scores at trajectories.hcompany.ai. For OSWorld that's 369 desktop tasks, each run 3 times, with the steps, tokens and time of every attempt. This role builds and runs the evaluation framework that produces runs like these.
H builds computer-use agents and the models behind them. Developers use them through a managed API, and our forward deployed engineers take them into enterprise workflows.
What this team ownsThe evaluation framework: orchestration, runtimes and observability. Researchers and forward deployed engineers bring the benchmarks, across web apps, desktop applications and the command line. Your job is to make the framework that runs them reliable, fast and cheap, and to make adding a new one quick. Research uses the results to choose checkpoints and decide whether a model ships. Product and the forward deployed engineers use them to measure agents on customer workflows. It carries roughly 50 benchmarks now. That number should be between 100 and 200 soon, and the framework has to keep up.
What you'd be doingIntegration support for researchers and forward deployed engineers bringing in a benchmark, with a shorter path each time.
Setting the standard for how a benchmark enters the framework, and building the checks that enforce it.
Scheduling and observability, so cluster capacity isn't left idle while evaluation jobs queue.
Reproducible results across trials, so a release decision rests on numbers that hold.
Whatever stack a benchmark calls for. One week that's cluster tuning; the next it's a browser extension or desktop environments.
Time with customers, from single developers to large companies, to