TheAgentCompany
Real workplace tasks across code, chat, documents, spreadsheets, and project tools. Every harness ran the same ten-task slice with GLM-5.3-Flash in freshly reset company environments.
Cost per full completion · same model and task slice
$0.0225 / completion · 26% less than Codex
Naive Managed Agents fully completed six of ten tasks and earned the highest mean reward in the field. It matched Stock Pi's completion count at 78% lower cost per completion and beat every other harness on both mean reward and completion efficiency.
Efficiency leaderboard?Published harness results on the same benchmark configuration, ordered by measured model cost per successful task.
Agent-model cost per fully completed task. Every published harness used the same tasks, model, provider rate card, and one-attempt protocol.
| Harness | Cost / successful task | Tasks solved | Success rate | Time / task | Input tokens / task | Output tokens / task | Turns / task |
|---|---|---|---|---|---|---|---|
| Naive Managed Agents | $0.0225 | 6 / 10 | 60.0% | 13.3 min | 643.5k | 9.1k | 30.5 |
| Codex | $0.0303 | 5 / 10 | 50.0% | 7.1 min | 794.0k | 2.9k | 1.0 |
| Claude Code | $0.0346 | 5 / 10 | 50.0% | 8.1 min | 765.0k | 7.4k | 1.0 |
| Hermes | $0.0385 | 5 / 10 | 50.0% | 14.1 min | 926.4k | 15.6k | 43.8 |
| Stock Pi | $0.102 | 6 / 10 | 60.0% | 7.3 min | 3.3M | 2.6k | 48.1 |
A full completion is a task with official reward 1.0. Mean reward is the official weighted partial-completion score, so it preserves useful progress on tasks that did not fully complete. This is a matched ten-task study, not the full benchmark leaderboard.
Why we run this benchmark
Coding benchmarks cover only one part of knowledge work. TheAgentCompany requires an agent to move between browser-based business systems, communication tools, office files, and repositories while preserving task intent across a multi-tool workflow.
Methodology
Five harnesses each attempted the same ten TheAgentCompany 1.0.0 task images once. Each attempt received the same prompt, credentials, browser CLI, GLM-5.3-Flash model, low reasoning setting, 90-minute limit, and fresh container, home, browser, and service state. All 50 trials were graded and none were excluded.
Reward and full completion come from the benchmark evaluator. Cost includes the tested harness's measured uncached input, cache-read input, and output tokens repriced on one shared rate card. The optimized E00 result is shown publicly as Naive Managed Agents; no additional Naive Managed Agents variants are included.
Want the winning configuration? Read the docs or see pricing.