naïve
← Benchmarks

Long Horizon Terminal Bench

Multi-hour terminal tasks run unattended in real sandboxes, graded by verification tests. Same model everywhere — glm-5.2 — only the harness is swapped, so every gap on this page is harness, not model.

Last run: August 2026Task suite: terminal-bench ↗

Tasks solved · vs Claude Code, same window

13 / 16 solved · 2.9× cheaper

Naive Managed Agents on the prioritywindow solved thirteen of sixteen tasks, the highest count in the field. At the like-for-like default window, each successful task cost 2.9× less than Claude Code ($0.2976 vs $0.8720), while Naive Managed Agents solved twelve tasks to Claude Code's eleven.

Efficiency leaderboard?All harness × window cells — filter by window, click a column to re-sort.

The completion window is a single field on the request — immediate answers now, priority soon, loose eventually. Only same-window rows are a like-for-like comparison.

?Show all harness × window cells, or only one completion window.

Sorted by Cost / successful task ↑

#HarnessWindow
1Naive Managed Agentsloose ◆$0.170910 / 1662.5%45.1 min867.0k16.2k38.5
2Naive Managed Agentspriority ◆$0.233213 / 1681.3%50.8 min1.2M20.5k40.1
3Claude Codeloose$0.293111 / 1668.8%58.4 min1.8M21.8k41.9
4Naive Managed Agentsimmediate$0.297612 / 1675.0%24.7 min743.1k18.0k31.3
5Hermesloose$0.34697 / 1546.7%56.0 min1.6M10.7k37.2
6Codexloose$0.383311 / 1670.2%54.0 min2.1M48.2k44.6
7Hermespriority$0.532711 / 1668.8%55.6 min2.6M15.4k48.9
8Claude Codepriority$0.572512 / 1675.0%62.7 min2.7M29.9k51.5
9Claude Codeimmediate$0.872011 / 1668.8%25.7 min2.1M21.7k47.4
10Codexpriority$1.01108 / 1651.4%56.8 min2.3M44.7k43.3
11Hermesimmediate$1.095010 / 1662.5%37.5 min1.9M11.9k41.8
12Piimmediate$1.10169 / 1656.3%29.6 min1.0M23.0k36.4
13Codeximmediate$1.35959 / 1657.7%26.2 min1.4M34.8k41.8
14Pipriority—6 / 1637.5%32.7 min578.0k18.6k30.1
15Piloose—6 / 1637.5%33.5 min758.5k18.8k30.6

◆ on the cost-efficiency frontier — no other configuration both solves more tasks and costs less per successful task. Dollars are what the vendor billed, not list price; cells without a verified invoice reading show no cost and are unranked on spend.

Why we run this benchmark

Most benchmarks measure a model; this one measures the machinery around it. Over multi-hour horizons the harness differences compound — context budgeting, checkpointing, recovery from failed builds. Running the same model through every harness across all three completion windows isolates the one variable we sell, with Naive Managed Agents benchmarked as just another row against the strongest public alternatives.

Methodology

Tasks come from the long-horizon split of terminal-bench, pinned to a fixed dataset commit. Each attempt starts from a clean containerized sandbox and runs unattended until it finishes or its window expires. Grading is binary and automatic — no partial credit, no human judging.

Every harness runs the same open-weights model — glm-5.2 — at identical sampling settings, with repeated attempts per task per window. Dollars are measured, not modelled: every figure is corrected against vendor invoices rather than list price, because rate-card accounting over-states spend by a different factor for every harness and would mis-rank them. Configurations whose invoice readings could not be verified carry no cost and are unranked on spend; attempts that failed for infrastructure reasons are excluded rather than estimated.

Want the winning configuration? Read the docs or see pricing.