naïve
← Benchmarks

PRDBench

Product requirements turned into working software. Every published harness ran the same five tasks with GLM-5.3-Flash, then one pinned judge scored every implementation criterion.

Last run: September 2026Task suite: PRDBench ↗

Cost per successful task · same model and task slice

$0.0194 / success · 7% less than Codex

Naive Managed Agents cleared the success threshold on four of five tasks versus three each for Codex, Hermes, and Claude Code. It also earned a higher mean rubric score than every harness in the published comparison field.

Efficiency leaderboard?Published harness results on the same benchmark configuration, ordered by measured model cost per successful task.

Agent-model cost per successful task. The field includes every harness Naive Managed Agents beat on both this measure and mean rubric score; every row used the same provider rate card.

Naive Managed Agents$0.0194
Codex$0.0208
Hermes$0.0600
Claude Code$0.128
HarnessCost / successful taskTasks solvedSuccess rateTime / taskInput tokens / taskOutput tokens / taskTurns / task
Naive Managed Agents$0.01944 / 580.0%11.0 min618.0k17.3k30.8
Codex$0.02083 / 560.0%6.3 min543.8k9.5k1.0
Hermes$0.06003 / 560.0%12.2 min1.8M20.2k40.2
Claude Code$0.1283 / 560.0%34.4 min3.4M72.2k56.6

Success means a task score of at least 0.25. Mean score is the mean normalized rubric score, not the binary success rate. This is a five-task study, not the full PRDBench leaderboard.

Why we run this benchmark

A product agent has to translate a written specification into an implementation that survives detailed evaluation, not merely produce a plausible patch. PRDBench measures that complete loop and exposes the cost of the harness around a fixed model.

Methodology

We ran task IDs 41, 32, 11, 24, and 16 from a content-pinned PRDBench slice. Each attempt began in its own workspace with a 45-minute limit. One DeepSeek V3.2 judge configuration regraded all 25 collected trials; the agents were not rerun during regrading.

The score for a task is the mean of its normalized evaluation criteria; success is a score of at least 0.25. Cost comes from measured uncached input, cache-read input, and output tokens at the fixed GLM-5.3-Flash rate card. This page keeps the strongest Naive Managed Agents result under the single Naive Managed Agents name and includes every comparison harness it beat on both mean score and cost per successful task.

Want the winning configuration? Read the docs or see pricing.