Business Bench
Practical business work across spreadsheets, bookkeeping, reports, document reformatting, extraction, drafting, and tooling. Every harness ran the same 25-task stratified slice with GLM-5.3-Flash.
Cost per successful task · same model and task slice
$0.0223 / success · tied for most tasks solved
Naive Managed Agents solved 12 of 25 tasks, tying Stock Pi and Hermes for the most successes in the field. At $0.0223 per successful task, it was 13% less expensive than Stock Pi and 68% less expensive than Hermes.
Efficiency leaderboard?Published harness results on the same benchmark configuration, ordered by measured model cost per successful task.
Agent-model cost per successful task. Every harness used the same tasks, model, provider route, rate card, and one-attempt protocol.
| Harness | Cost / successful task | Tasks solved | Success rate | Time / task | Input tokens / task | Output tokens / task | Turns / task |
|---|---|---|---|---|---|---|---|
| Naive Managed Agents | $0.0223 | 12 / 25 | 48.0% | 3.8 min | 141.6k | 7.5k | 12.0 |
| Claude Code | $0.0518 | 10 / 25 | 40.0% | 3.4 min | 320.4k | 7.7k | 11.9 |
| Codex | $0.294 | 7 / 25 | 28.0% | 8.2 min | 1.8M | 15.2k | 37.9 |
| Hermes | $0.0688 | 12 / 25 | 48.0% | 5.6 min | 533.8k | 9.5k | 17.7 |
| Stock Pi | $0.0258 | 12 / 25 | 48.0% | 6.3 min | 124.9k | 10.3k | 12.2 |
A task succeeds only when every required official check passes. Benchmark failures and task-limit timeouts remain valid unsuccessful attempts. This is a matched 25-task study, not the full benchmark leaderboard.
Why we run this benchmark
Coding tests cover only one part of day-to-day work. Business Bench measures whether an agent can turn messy source files and written instructions into correct spreadsheets, reports, reconciliations, notices, and other concrete business deliverables.
Methodology
Five harnesses each attempted the same stratified sample of 25 Business Bench tasks once. The sample used seed 20260928 and spans seven task categories. Each attempt ran in the official task container with an isolated workspace and home directory, using GLM-5.3-Flash through OpenRouter at low reasoning effort.
The benchmark's deterministic checks grade the produced files. Infrastructure-invalid attempts were recovered before publication; ordinary harness failures and eight task-limit timeouts were kept. All 125 published trials are graded and none are excluded.
Cost is calculated from each harness's recorded uncached input, cache-read input, and output tokens on one pinned OpenRouter Z.AI rate card. Time, token, and turn figures are means across all 25 attempts, including unsuccessful ones.
Want the winning configuration? Read the docs or see pricing.