SWE-Bench Pro
Real repository issues resolved end-to-end, graded by hidden tests. Naive Managed Agents, Codex, Claude Code, Hermes, and Stock Pi each ran the same fifteen-task slice with GLM-5.3-Flash.
Cost per successful task · same model and task slice
$0.0398 / success · 13% cheaper
Naive Managed Agents solved ten of fifteen tasks at $0.0398 per successful task. The next-cheapest harness, Codex, solved eight at $0.0460 each. Hermes solved twelve at $0.2510 each, Claude Code solved all fifteen at $2.7337 each, and Stock Pi solved nine at $2.6998 each.
Efficiency leaderboard?Five matched harnesses on one model and task slice; click a column to re-sort.
Every row uses the same model, provider route, task slice, and one-attempt protocol. The harness determines how much context, time, and model spend each attempted task consumes.
Sorted by Cost / successful task ↑
| # | Harness | Model | |||||||
|---|---|---|---|---|---|---|---|---|---|
| 1 | Naive Managed Agents | GLM-5.3-Flash ◆ | $0.0398 | 10 / 15 | 66.7% | 10.1 min | 811.7k | 8.6k | 38.9 |
| 2 | Codex | GLM-5.3-Flash | $0.0460 | 8 / 15 | 53.3% | 6.6 min | 1.3M | 9.0k | — |
| 3 | Hermes | GLM-5.3-Flash ◆ | $0.2510 | 12 / 15 | 80.0% | 32.6 min | 2.6M | 13.7k | — |
| 4 | Pi | GLM-5.3-Flash | $2.6998 | 9 / 15 | 60.0% | 19.7 min | 13.6M | 22.9k | — |
| 5 | Claude Code | GLM-5.3-Flash ◆ | $2.7337 | 15 / 15 | 100.0% | 15.9 min | 4.0M | 26.1k | — |
◆ on the cost-efficiency frontier — no other configuration both solves more tasks and costs less per successful task. Every row uses GLM-5.3-Flash on the same task slice.
Why we run this benchmark
Issue resolution is the workload most teams actually buy an agent for. Holding the model and tasks fixed isolates how much completion rate, latency, and spend come from the machinery around the model, with Naive Managed Agents benchmarked as another row against public alternatives.
Methodology
Tasks come from SWE-Bench Pro, pinned to a fixed dataset commit. Each attempt starts from a clean containerized checkout of the repository at the issue commit and runs unattended; grading is the suite's hidden tests — binary and automatic, no partial credit, no human judging.
Every harness ran the identical fifteen-task slice with GLM-5.3-Flash through OpenRouter. Invalid infrastructure attempts were repaired and rerun rather than scored as task failures. The published field contains 75 usable trials.
Cost is the measured model API charge recorded by each run. Hermes did not emit dollar totals, so its recorded input and output tokens are repriced on the shared GLM-5.3-Flash rate card; its missing cache split is disclosed in the data notes.
Want the winning configuration? Read the docs or see pricing.