YC-Bench
A long-horizon startup simulation with hiring, research, payroll, deadlines, and bankruptcy risk. Five harnesses attempted the same year-long scenario on three seeds with GLM-5.3-Flash; incomplete and bankrupt outcomes remain visible.
Mean cost per completed year · three matched seeds
$0.825 / success · 29% less
Naive Managed Agents completed all three simulated years for 29% less model spend than Stock Pi. It also finished with higher mean funds ($1.93M vs $1.70M) and more successfully completed work.
Efficiency leaderboard?Published harness results on the same benchmark configuration, ordered by measured model cost per successful task.
Measured model cost per successfully completed simulated year. Every harness used the same model, three-seed task set, and price card.
| Harness | Cost / successful task | Tasks solved | Success rate | Time / task | Input tokens / task | Output tokens / task | Turns / task |
|---|---|---|---|---|---|---|---|
| Naive Managed Agents | $0.825 | 3 / 3 | 100.0% | 4.59 h | 26.4M | 587.7k | 161.7 |
| Codex | $0.357 | 2 / 3 | 66.7% | 1.79 h | 8.7M | 135.6k | 70.7 |
| Stock Pi | $1.156 | 3 / 3 | 100.0% | 2.55 h | 40.0M | 205.8k | 144.7 |
| Exo | — | 0 / 3 | 0.0% | 8.93 h* | 32.2M | 919.3k | 67.0 |
| Claude Code | $5.483 | 2 / 3 | 66.7% | 12.50 h | 148.1M | 1.8M | 141.3 |
Final funds and completed jobs are terminal-outcome means. Codex and Claude Code each reached two full years from three attempts; Codex's third run went bankrupt, and Claude's third stopped late in the year. Exo values marked * come from its single substantive partial run; its other two attempts failed before the simulation started.
Why we run this benchmark
Short coding tasks do not test whether an agent can preserve a plan through delayed feedback. YC-Bench makes the agent allocate capital, manage staff, and react to a year of changing constraints without a human steering each turn.
Methodology
Three deterministic seeds (1–3) used the canonical one-year scenario. Each harness controlled the same startup through the benchmark CLI until the horizon ended, bankruptcy occurred, or the harness stopped. We read final funds, task outcomes, and terminal reason from every substantive benchmark artifact.
Cost is reconstructed from each harness trace at the pinned GLM-5.3-Flash rate card: uncached input, cache-read input, and output tokens are priced separately. Means include all substantive terminal artifacts, including partial and bankrupt runs; zero-turn startup failures are counted as attempts but excluded from token, time, funds, and job averages.
Want the winning configuration? Read the docs or see pricing.