Long Horizon Terminal Bench
Multi-hour terminal tasks run unattended in real sandboxes, graded by verification tests. Same model everywhere — glm-5.2 — only the harness is swapped, so every gap on this page is harness, not model.
Tasks solved · vs Claude Code, same window
13 / 16 solved · 2.9× cheaper
Naive Managed Agents on the prioritywindow solved thirteen of sixteen tasks, the highest count in the field. At the like-for-like default window, each successful task cost 2.9× less than Claude Code ($0.2976 vs $0.8720), while Naive Managed Agents solved twelve tasks to Claude Code's eleven.
Efficiency leaderboard?All harness × window cells — filter by window, click a column to re-sort.
The completion window is a single field on the request — immediate answers now, priority soon, loose eventually. Only same-window rows are a like-for-like comparison.
Sorted by Cost / successful task ↑
| # | Harness | Window | |||||||
|---|---|---|---|---|---|---|---|---|---|
| 1 | Naive Managed Agents | loose ◆ | $0.1709 | 10 / 16 | 62.5% | 45.1 min | 867.0k | 16.2k | 38.5 |
| 2 | Naive Managed Agents | priority ◆ | $0.2332 | 13 / 16 | 81.3% | 50.8 min | 1.2M | 20.5k | 40.1 |
| 3 | Claude Code | loose | $0.2931 | 11 / 16 | 68.8% | 58.4 min | 1.8M | 21.8k | 41.9 |
| 4 | Naive Managed Agents | immediate | $0.2976 | 12 / 16 | 75.0% | 24.7 min | 743.1k | 18.0k | 31.3 |
| 5 | Hermes | loose | $0.3469 | 7 / 15 | 46.7% | 56.0 min | 1.6M | 10.7k | 37.2 |
| 6 | Codex | loose | $0.3833 | 11 / 16 | 70.2% | 54.0 min | 2.1M | 48.2k | 44.6 |
| 7 | Hermes | priority | $0.5327 | 11 / 16 | 68.8% | 55.6 min | 2.6M | 15.4k | 48.9 |
| 8 | Claude Code | priority | $0.5725 | 12 / 16 | 75.0% | 62.7 min | 2.7M | 29.9k | 51.5 |
| 9 | Claude Code | immediate | $0.8720 | 11 / 16 | 68.8% | 25.7 min | 2.1M | 21.7k | 47.4 |
| 10 | Codex | priority | $1.0110 | 8 / 16 | 51.4% | 56.8 min | 2.3M | 44.7k | 43.3 |
| 11 | Hermes | immediate | $1.0950 | 10 / 16 | 62.5% | 37.5 min | 1.9M | 11.9k | 41.8 |
| 12 | Pi | immediate | $1.1016 | 9 / 16 | 56.3% | 29.6 min | 1.0M | 23.0k | 36.4 |
| 13 | Codex | immediate | $1.3595 | 9 / 16 | 57.7% | 26.2 min | 1.4M | 34.8k | 41.8 |
| 14 | Pi | priority | — | 6 / 16 | 37.5% | 32.7 min | 578.0k | 18.6k | 30.1 |
| 15 | Pi | loose | — | 6 / 16 | 37.5% | 33.5 min | 758.5k | 18.8k | 30.6 |
◆ on the cost-efficiency frontier — no other configuration both solves more tasks and costs less per successful task. Dollars are what the vendor billed, not list price; cells without a verified invoice reading show no cost and are unranked on spend.
Why we run this benchmark
Most benchmarks measure a model; this one measures the machinery around it. Over multi-hour horizons the harness differences compound — context budgeting, checkpointing, recovery from failed builds. Running the same model through every harness across all three completion windows isolates the one variable we sell, with Naive Managed Agents benchmarked as just another row against the strongest public alternatives.
Methodology
Tasks come from the long-horizon split of terminal-bench, pinned to a fixed dataset commit. Each attempt starts from a clean containerized sandbox and runs unattended until it finishes or its window expires. Grading is binary and automatic — no partial credit, no human judging.
Every harness runs the same open-weights model — glm-5.2 — at identical sampling settings, with repeated attempts per task per window. Dollars are measured, not modelled: every figure is corrected against vendor invoices rather than list price, because rate-card accounting over-states spend by a different factor for every harness and would mis-rank them. Configurations whose invoice readings could not be verified carry no cost and are unranked on spend; attempts that failed for infrastructure reasons are excluded rather than estimated.
Want the winning configuration? Read the docs or see pricing.