CoffeeBench
Agents operated one roaster inside a six-company coffee economy: sourcing inventory, negotiating, roasting, selling, and paying bills. Every company used GLM-5.3-Flash.
Cost per successful task · five matched runs
$0.337 / success · 58% less
Naive Managed Agents completed all five 14-day simulations at 58% less focal-agent spend than Stock Pi. The result uses isolated, request-level model ledgers.
Efficiency leaderboard?Published harness results on the same benchmark configuration, ordered by measured model cost per successful task.
Focal-agent model cost per completed 14-day simulation. The other five simulator agents are shared background actors and excluded from this comparison.
| Harness | Cost / successful task | Tasks solved | Success rate | Time / task | Input tokens / task | Output tokens / task | Turns / task |
|---|---|---|---|---|---|---|---|
| Naive Managed Agents | $0.337 | 5 / 5 | 100.0% | 1.72 h | 10.5M | 171.6k | 693.6 |
| Stock Pi | $0.797 | 5 / 5 | 100.0% | 2.47 h | 21.4M | 189.0k | 722.8 |
| Claude Code | $1.514 | 5 / 5 | 100.0% | 18.71 h | 88.6M | 2.1M | — |
| Exo | $2.344 | 5 / 5 | 100.0% | 4.54 h | 148.4M | 138.6k | 557.8 |
This is a modified 14-day slice, not CoffeeBench's canonical 90-day score. Net income is descriptive: the Naive Managed Agents result used seeds 10–14, while reference arms used seeds 0–4, so small outcome differences are not paired estimates.
Why we run this benchmark
CoffeeBench adds a live economy and other autonomous actors. Useful behavior depends on repeated observation, negotiation, inventory timing, and cash discipline, making it a demanding test of whether lower inference cost survives an extended business workflow.
Methodology
Each run controlled roaster_Afor 14 simulated days while five fixed background agents used the same model. The published cost is the tested harness's focal-agent spend; business results come from CoffeeBench's accounting ledger.
Naive Managed Agents' five runs use fresh per-decision sessions, medium effort, verification disabled, and audited request ledgers. Reference-arm aggregates were cross-checked against the benchmark database. When multiple Naive Managed Agents configurations were available, this page keeps only the strongest Naive Managed Agents result under the single Naive Managed Agents name.
Want the winning configuration? Read the docs or see pricing.