naïve
← Benchmarks

TheAgentCompany

Real workplace tasks across code, chat, documents, spreadsheets, and project tools. Every harness ran the same ten-task slice with GLM-5.3-Flash in freshly reset company environments.

Last run: September 2026Task suite: TheAgentCompany ↗

Cost per full completion · same model and task slice

$0.0225 / completion · 26% less than Codex

Naive Managed Agents fully completed six of ten tasks and earned the highest mean reward in the field. It matched Stock Pi's completion count at 78% lower cost per completion and beat every other harness on both mean reward and completion efficiency.

Efficiency leaderboard?Published harness results on the same benchmark configuration, ordered by measured model cost per successful task.

Agent-model cost per fully completed task. Every published harness used the same tasks, model, provider rate card, and one-attempt protocol.

Naive Managed Agents$0.0225
Codex$0.0303
Claude Code$0.0346
Hermes$0.0385
Stock Pi$0.102
HarnessCost / successful taskTasks solvedSuccess rateTime / taskInput tokens / taskOutput tokens / taskTurns / task
Naive Managed Agents$0.02256 / 1060.0%13.3 min643.5k9.1k30.5
Codex$0.03035 / 1050.0%7.1 min794.0k2.9k1.0
Claude Code$0.03465 / 1050.0%8.1 min765.0k7.4k1.0
Hermes$0.03855 / 1050.0%14.1 min926.4k15.6k43.8
Stock Pi$0.1026 / 1060.0%7.3 min3.3M2.6k48.1

A full completion is a task with official reward 1.0. Mean reward is the official weighted partial-completion score, so it preserves useful progress on tasks that did not fully complete. This is a matched ten-task study, not the full benchmark leaderboard.

Why we run this benchmark

Coding benchmarks cover only one part of knowledge work. TheAgentCompany requires an agent to move between browser-based business systems, communication tools, office files, and repositories while preserving task intent across a multi-tool workflow.

Methodology

Five harnesses each attempted the same ten TheAgentCompany 1.0.0 task images once. Each attempt received the same prompt, credentials, browser CLI, GLM-5.3-Flash model, low reasoning setting, 90-minute limit, and fresh container, home, browser, and service state. All 50 trials were graded and none were excluded.

Reward and full completion come from the benchmark evaluator. Cost includes the tested harness's measured uncached input, cache-read input, and output tokens repriced on one shared rate card. The optimized E00 result is shown publicly as Naive Managed Agents; no additional Naive Managed Agents variants are included.

Want the winning configuration? Read the docs or see pricing.