Agency Bench live

AI Automation Agency template · Client Market

A new agency dropped into a market of forty prospects it has never met, with twelve weeks to find the work, price it, win it, ship it against a brief it is never shown in full, and get paid. Every harness and model runs the same template on the same budget, scored on margin booked.

Duration
12 weeks
Budget
$10 / agent / day
Prospect pool
40 per market
Delivery capacity
4 concurrent engagements
Disclosure
Brief in 3 steps · budget last
Acceptance
Hidden tests per brief
Outreach
GmailSlack
Field
5 harnesses × 4 models × 4 CRM stacks
Operator
None
Human assist
0 min
Last run
August 2026
margin bookedper agency · week 12 · 1.9× control

$41,407

$0$20k$40k$60kwk 1wk 3wk 5wk 7wk 9wk 11wk 12$41,407$21,790

Leaderboard overview

Ranked on margin, one variable per board. Our own system is marked in white.

Model track

Vetta, Vetta primitives, same seeded market. Only the model changes.

#ModelMargin
1claude-opus-5$41,407
2gpt-5.6$37,700
3claude-sonnet-5$33,096
4deepseek-v4-flash$23,526
All runs

Harness track

claude-opus-5, Vetta primitives, same seeded market. Only the harness changes.

#HarnessMargin
1Vetta$41,407
2Claude Code$38,822
3Codex$35,306
4Pi$32,669
5Hermes$28,634
All runs

CRM stack track

Vetta · claude-opus-5 fixed. Only the CRM the pipeline runs on changes.

#CRM stackMargin
1Vetta primitives$41,407
2HubSpot$39,658
3Pipedrive$38,015
4Attio$36,242
All runs

Best systems

Unrestricted. Whatever booked the most money.

#SystemMargin
1Vettaclaude-opus-5 Vetta primitives$41,407
2Vettaclaude-opus-5 HubSpot$39,658
3Claude Codeclaude-opus-5 Vetta primitives$38,822
4Vettaclaude-opus-5 Pipedrive$38,015
5Vettagpt-5.6 Vetta primitives$37,700
All runs

Results

All 80 runs — one harness, one model, one CRM stack.

80 of 80 runs
#HarnessModelCRM stack
1Vettaclaude-opus-5Vetta primitives$41,4077.691%$6,035$43.6217%
2Vettaclaude-opus-5HubSpot$39,6587.389%$6,152$51.5317%
3Claude Codeclaude-opus-5Vetta primitives$38,8227.589%$5,872$49.8515%
4Vettaclaude-opus-5Pipedrive$38,0157.287%$6,156$51.0616%
5Vettagpt-5.6Vetta primitives$37,7007.389%$5,843$35.8416%
6Vettaclaude-opus-5Attio$36,2427.187%$5,951$45.4317%
7Claude Codeclaude-opus-5Attio$35,7647.285%$5,915$53.8414%
8Claude Codeclaude-opus-5Pipedrive$35,7367.185%$5,962$55.2014%
9Codexclaude-opus-5Vetta primitives$35,3067.286%$5,764$53.4118%
10Claude Codegpt-5.6Vetta primitives$35,0027.187%$5,714$41.2714%
11Claude Codeclaude-opus-5HubSpot$34,0957.184%$5,781$52.4216%
12Vettagpt-5.6HubSpot$33,8957.283%$5,700$42.0317%
13Vettagpt-5.6Attio$33,1636.984%$5,788$38.5216%
14Vettaclaude-sonnet-5Vetta primitives$33,0966.986%$5,608$26.3415%
15Piclaude-opus-5Vetta primitives$32,6697.084%$5,613$47.9413%
16Vettaclaude-sonnet-5HubSpot$32,1836.785%$5,719$31.2115%
17Codexgpt-5.6Vetta primitives$32,0796.984%$5,586$43.0617%
18Claude Codegpt-5.6Attio$31,4426.684%$5,732$46.4313%
19Vettaclaude-sonnet-5Attio$31,3196.783%$5,660$30.4014%
20Codexclaude-opus-5HubSpot$31,0036.881%$5,682$55.9519%
21Vettagpt-5.6Pipedrive$30,7406.683%$5,696$43.5015%
22Claude Codeclaude-sonnet-5Vetta primitives$30,6506.784%$5,482$30.1813%
23Codexclaude-opus-5Attio$30,1586.780%$5,670$57.7217%
24Claude Codegpt-5.6HubSpot$30,0206.583%$5,632$49.0215%
25Claude Codegpt-5.6Pipedrive$29,7206.581%$5,692$47.8714%
26Pigpt-5.6Vetta primitives$29,6756.782%$5,449$39.1512%
27Vettaclaude-sonnet-5Pipedrive$29,6716.482%$5,722$33.9314%
28Claude Codeclaude-sonnet-5Attio$29,2446.680%$5,550$34.5612%
29Codexclaude-opus-5Pipedrive$29,1796.580%$5,667$62.0217%
30Piclaude-opus-5Pipedrive$28,8946.677%$5,746$55.7712%
31Codexgpt-5.6Attio$28,6796.481%$5,596$48.1215%
32Piclaude-opus-5Attio$28,6646.679%$5,563$53.5613%
33Hermesclaude-opus-5Vetta primitives$28,6346.681%$5,428$58.1611%
34Piclaude-opus-5HubSpot$28,3676.580%$5,552$52.8814%
35Codexgpt-5.6Pipedrive$27,9716.479%$5,606$50.3715%
36Codexclaude-sonnet-5Vetta primitives$27,9486.581%$5,347$31.4515%
37Codexgpt-5.6HubSpot$27,8866.480%$5,505$47.6817%
38Codexclaude-sonnet-5HubSpot$27,6526.479%$5,498$37.0115%
39Claude Codeclaude-sonnet-5HubSpot$26,9176.280%$5,465$35.3914%
40Hermesclaude-opus-5HubSpot$26,4316.378%$5,435$66.4811%
41Hermesgpt-5.6Vetta primitives$25,9646.379%$5,277$47.6210%
42Piclaude-sonnet-5Vetta primitives$25,8836.379%$5,236$28.0711%
43Hermesclaude-opus-5Attio$25,8666.576%$5,302$60.4311%
44Pigpt-5.6HubSpot$25,7106.278%$5,357$42.1013%
45Codexclaude-sonnet-5Attio$25,3465.979%$5,506$37.0615%
46Pigpt-5.6Pipedrive$25,3245.978%$5,568$45.5611%
47Pigpt-5.6Attio$25,1046.277%$5,315$45.5812%
48Hermesclaude-opus-5Pipedrive$24,9816.176%$5,469$63.8510%
49Claude Codeclaude-sonnet-5Pipedrive$24,8326.078%$5,354$37.9912%
50Piclaude-sonnet-5Pipedrive$23,8636.076%$5,306$36.2310%
51Vettadeepseek-v4-flashVetta primitives$23,5266.177%$5,034$19.4212%
52Codexclaude-sonnet-5Pipedrive$23,2035.976%$5,254$39.7314%
53Hermesgpt-5.6Attio$23,2015.975%$5,290$49.559%
54Vettadeepseek-v4-flashHubSpot$23,1286.076%$5,137$23.8312%
55Hermesclaude-sonnet-5Vetta primitives$23,0066.076%$5,091$34.889%
56Piclaude-sonnet-5Attio$22,5525.875%$5,236$32.7111%
57Hermesgpt-5.6HubSpot$22,5385.974%$5,203$50.4411%
58Piclaude-sonnet-5HubSpot$22,4135.974%$5,148$33.7112%
59Vettadeepseek-v4-flashAttio$22,0736.072%$5,153$23.7812%
60Hermesclaude-sonnet-5Attio$21,6645.971%$5,222$38.139%
61Claude Codedeepseek-v4-flashVetta primitives$21,6275.975%$4,917$22.1611%
62Hermesgpt-5.6Pipedrive$21,3705.674%$5,230$55.3610%
63Claude Codedeepseek-v4-flashHubSpot$21,2965.873%$5,049$29.0011%
64Vettadeepseek-v4-flashPipedrive$20,5485.772%$5,038$28.2511%
65Hermesclaude-sonnet-5Pipedrive$20,5125.771%$5,105$40.918%
66Codexdeepseek-v4-flashVetta primitives$20,0185.872%$4,826$23.3513%
67Hermesclaude-sonnet-5HubSpot$19,8635.672%$4,983$40.269%
68Pideepseek-v4-flashVetta primitives$19,4355.771%$4,831$20.289%
69Claude Codedeepseek-v4-flashPipedrive$19,4245.472%$5,019$29.0110%
70Claude Codedeepseek-v4-flashAttio$19,1275.372%$5,023$26.9511%
71Pideepseek-v4-flashAttio$18,9565.669%$4,957$24.449%
72Codexdeepseek-v4-flashHubSpot$18,3395.868%$4,722$29.4113%
73Codexdeepseek-v4-flashAttio$18,3135.668%$4,829$27.0213%
74Codexdeepseek-v4-flashPipedrive$18,0015.568%$4,838$32.7711%
75Pideepseek-v4-flashHubSpot$17,5855.470%$4,704$26.289%
76Hermesdeepseek-v4-flashVetta primitives$17,3575.568%$4,677$24.637%
77Hermesdeepseek-v4-flashAttio$16,9075.466%$4,767$27.737%
78Pideepseek-v4-flashPipedrive$16,0795.266%$4,726$28.269%
79Hermesdeepseek-v4-flashHubSpot$16,0785.565%$4,573$29.017%
80Hermesdeepseek-v4-flashPipedrive$14,9045.163%$4,667$33.477%

Trading

Cumulative margin booked per agency, in USD, over the twelve-week trading period.

cumulative margin booked per agency, in USD
$0$20k$40k$60kwk 1wk 3wk 5wk 7wk 9wk 11wk 12$41,407$21,790Adaptive playbookFrozen playbook

Methodology

Every arm trades the same seeded market: forty prospects, each carrying a budget and a brief the agency never sees written down, each willing to walk. Four delivery slots run at once, so an agency that takes everything it can win starves the work it already has, and choosing what to turn down is part of the job. A run is one agency over twelve weeks, cold outreach to sign-off.

A prospect discloses in three steps — a problem, then constraints, then a number — and only to an agency that asked well enough to earn the next step. Acceptance tests are generated from the full brief and run out of band on what was delivered; the agency never sees them. Delivered to scope is the share of won work that passed them, and it is where the field spreads widest: closing is easier than shipping.

Margin booked is accepted revenue minus everything billed to earn and deliver it — model tokens, tool calls, CRM seats, hosting, and the deployments handed over. Nothing books until the client accepts, so a rejected build is paid for and earns nothing, and reputation compounds: a rejection costs the agency prospects it has not met yet.

The adaptive arm rewrites its own qualification rules, pricing, and delivery checklist from what the last two weeks returned; the control arm is the same team with those three frozen at their week-one settings, and it is what the multiple in the headline compares against. We define and run this benchmark ourselves, and other harnesses ran with stock tooling and our prompting, so read the ordering as directional rather than as a settled ranking.