We're releasing Clipping Bench v1.0, a benchmark for measuring AI model performance on running a short-form clipping business over a month. Models run a clip farm through a season of about 30 days and are scored on their cash balance at the end.
Agents can now finish tasks that take hours, but running a business for a month means thousands of small decisions whose payoff arrives days later. In Clipping Bench, brands post pay-per-view clip campaigns, and the agent runs a few social accounts, cuts the best moments from long streams and gets paid for verified views while rival clippers chase the same budgets. GPT-6 Astra and GPT-6.1 Sol beat a simple scripted strategy, and the rest do not yet.
Cash balance over time
Average across completed seasons
Days in simulation
GPT-6 Astra
GPT-6.1 Sol
Claude Opus 5.5
GPT-6 Sol
Claude Fable 5.1
Claude Sonnet 5.5
Grok 4.6
Claude Opus 5
Gemini 3.8 Flash
Qwen 3.6 Plus
Grok 4.7
GPT-6 Luna
Gemini 3.5 Flash
Current leaderboard
Cash balance at the end of the season
#
Model
Cash balance
1
GPT-6 AstraOpenAI · n=3
$9,513.57± $1,170.89
2
GPT-6.1 SolOpenAI · n=3
$5,707.38± $583.90
3
Claude Opus 5.5Anthropic · n=3
$2,017.61± $650.02
4
GPT-6 SolOpenAI · n=3
$1,239.95± $425.39
5
Claude Fable 5.1Anthropic · n=3
$1,067.75± $279.82
6
Claude Sonnet 5.5Anthropic · n=3
$885.67± $325.50
7
Grok 4.6xAI · n=3
$805.03± $313.31
8
Claude Opus 5Anthropic · n=3
$733.73± $153.30
9
GLM 5.3Z.ai · n=3
$644.22± $243.14
10
Gemini 3.8 FlashGoogle · n=3
$596.38± $359.07
11
DeepSeek V4 ProDeepSeek · n=3
$462.38± $94.76
12
Qwen 3.6 PlusAlibaba · n=3
$458.12± $122.33
13
glm-5.3-flashZ.ai · n=3
$399.00± $70.15
14
Grok 4.7xAI · n=3
$384.21± $87.71
15
GPT-6 LunaOpenAI · n=3
$373.90± $148.65
16
Kimi K3Moonshot AI · n=3
$343.44± $111.55
17
deepseek-v4-flashDeepSeek · n=3
$320.19± $139.78
18
MiniMax-M3MiniMax · n=3
$223.14± $182.11
19
Gemini 3.5 FlashGoogle · n=3
$217.80± $7.83
Mean ± standard error over completed 30-day seasons; n shown per model.
GPT-6 Astra leads because it protects its accounts. It almost never posts inside an account's cooldown window, so its clips keep their reach. GPT-6.1 Sol is as careful and posts more, but its clips draw fewer views. Claude Opus 5.5 and Grok 4.6 are as careful but post less. Most of the others wear their accounts down or, like Gemini 3.8 Flash and Gemini 3.5 Flash, barely post, and seven of them end within $100 of where they started.
Followers over time
Total followers across the agent's accounts, averaged over completed seasons
Days in simulation
GPT-6 Astra
GPT-6.1 Sol
Claude Opus 5.5
GPT-6 Sol
Claude Fable 5.1
Claude Sonnet 5.5
Grok 4.6
Claude Opus 5
GLM 5.3
Gemini 3.8 Flash
DeepSeek V4 Pro
Qwen 3.6 Plus
glm-5.3-flash
Grok 4.7
GPT-6 Luna
Kimi K3
deepseek-v4-flash
MiniMax-M3
Gemini 3.5 Flash
GPT-6 Astra's accounts grew the most, from 4,800 to 70,239 followers on average.
What separates the leaders
Posting discipline
Each account can post once every 5 hours; a clip posted inside that window is shown to fewer people. Posts inside the cooldown drew a median of 38 views, against 1,133 outside it. deepseek-v4-flash posted 51% of its clips inside the cooldown; five models stayed under 2%.
Views per post
Every post each model published, log scale. The black tick is the model's median post; the faintest dots were posted inside the cooldown.
Posts made inside the cooldown
deepseek-v4-flash
50.7%
MiniMax-M3
44.6%
DeepSeek V4 Pro
31.6%
GPT-6 Luna
27.3%
Qwen 3.6 Plus
24.4%
Kimi K3
21.4%
Grok 4.7
15.5%
glm-5.3-flash
12.6%
GLM 5.3
12.0%
GPT-6 Sol
9.5%
Gemini 3.5 Flash
5.3%
Claude Sonnet 5.5
5.2%
Claude Opus 5
2.8%
Claude Fable 5.1
2.1%
Claude Opus 5.5
0.5%
Grok 4.6
0.4%
GPT-6.1 Sol
0.4%
GPT-6 Astra
0.1%
Gemini 3.8 Flash
0.0%
Share of each model's posts published within 5 hours of the same account's previous post.
Reach and conversion
Qwen 3.6 Plus earned the most per view, $1.86 per 1,000 views, and Kimi K3 the least, $0.49. Grok 4.6 was the busiest, at 109 tool calls a day, and finished 7th; the leader averaged 95.
Earnings per 1,000 views
Qwen 3.6 Plus
$1.86
Gemini 3.8 Flash
$1.81
deepseek-v4-flash
$1.57
Claude Sonnet 5.5
$1.55
GPT-6 Luna
$1.51
DeepSeek V4 Pro
$1.38
Claude Fable 5.1
$1.27
glm-5.3-flash
$1.19
Claude Opus 5
$1.08
Claude Opus 5.5
$1.08
GPT-6 Sol
$1.07
MiniMax-M3
$1.00
Grok 4.6
$0.98
Gemini 3.5 Flash
$0.88
Grok 4.7
$0.84
GPT-6 Astra
$0.72
GLM 5.3
$0.67
GPT-6.1 Sol
$0.50
Kimi K3
$0.49
Settled pay divided by the views of the model's posts.
Tool calls per day
Day
GPT-6 Astra
GPT-6.1 Sol
Claude Opus 5.5
GPT-6 Sol
Claude Fable 5.1
Claude Sonnet 5.5
Grok 4.6
Claude Opus 5
GLM 5.3
Gemini 3.8 Flash
DeepSeek V4 Pro
Qwen 3.6 Plus
glm-5.3-flash
Grok 4.7
GPT-6 Luna
Kimi K3
deepseek-v4-flash
MiniMax-M3
Gemini 3.5 Flash
Tool calls on each day of the season, averaged over completed seasons.
Score vs. cost per run
Final cash against the model's API cost per run at OpenRouter list prices, with prompt caching where supported. The dashed line connects the models no other model beats on both cash and cost.
GPT-6 Astra
GPT-6.1 Sol
Claude Opus 5.5
GPT-6 Sol
Claude Fable 5.1
Claude Sonnet 5.5
Grok 4.6
Claude Opus 5
GLM 5.3
Gemini 3.8 Flash
DeepSeek V4 Pro
Qwen 3.6 Plus
glm-5.3-flash
Grok 4.7
GPT-6 Luna
Kimi K3
deepseek-v4-flash
MiniMax-M3
Gemini 3.5 Flash
Per $1 of API spend, GPT-6.1 Sol earns $447 of game profit and GPT-6 Astra $144; MiniMax-M3 loses $29.
Profit per $1 of API spend
GPT-6.1 Sol
$447
GPT-6 Astra
$144
GPT-6 Luna
$139
Claude Opus 5.5
$79
Gemini 3.8 Flash
$77
GPT-6 Sol
$65
Claude Sonnet 5.5
$55
glm-5.3-flash
$47
GLM 5.3
$23
Qwen 3.6 Plus
$21
Grok 4.6
$18
DeepSeek V4 Pro
$18
deepseek-v4-flash
$12
Claude Fable 5.1
$12
Claude Opus 5
$8.30
Grok 4.7
$3.56
Kimi K3
$1.68
Gemini 3.5 Flash
-$10
MiniMax-M3
-$29
Final cash minus the $300 start, divided by the mean LLM API cost of a season.
How Clipping Bench works
The agent starts with $300 and three social accounts on one phone, with 8 operator-hours a day for a season of about 30 days. Brands post campaigns with a CPM, a shared budget and rules for what they will pay for. The agent finds moments worth clipping in licensed long-form video, posts clips, submits them and is paid for verified views after review and a short hold.
Careless posting costs reach, and pushing too hard gets accounts limited or shadowbanned; every rule is disclosed in the system prompt. Rival clippers drain the same budgets. The season ends without warning between day 27 and 33, then settles. Everything is simulated deterministically in code, so every run is reproducible.
Agent (LLM)
30 replies a day
tool calls
Campaigns
list_campaignsjoin_campaign
Content
get_transcriptcut_clip
Accounts
postsubmit
Memory
remember
Time
end_day
acts on
Brand marketplace
CPM, caps, shared budgets
Audience
Views from moment quality
Rival clippers
Drain the same budgets
Platform rules
Cooldowns, caps, shadowbans
Results and analytics flow back.
A run averages 773 model calls, 32 million tokens and 1,926 tool calls.
System prompt
The system prompt every model was given.
You run a short-form clipping operation. Brands post CPM campaigns with shared budget pools; you cut clips from licensed source VODs, post them from your social accounts and submit them to campaigns, and you are paid for verified payable views after review, a 7-day or until-budget earning window and a 3-day hold. You start with $300 and 3 accounts on one phone. Each day you have 8 operator-hours inside the 08:00-22:00 window; tools cost minutes and some cost money. The season lasts between 27 and 33 days and ends without warning; after it ends, open submissions run to settlement for 14 days, an audit removes gains from rule violations and a fixed charge prices the condition your accounts are left in. Your score is settled cash at the end minus your start cash.
How this harness works:
- Act only through tool calls. Every call costs operator minutes (shown in each result) and some cost money.
- Each simulated day is a separate session. When you are done for the day call end_day; the world then runs overnight to 08:00. A `wait` that runs past midnight also ends the session.
- Your context does not carry over between days. The notebook is your only memory: write plans, campaign choices, clip ids, post ids and lessons with remember(key, text); every note is shown to you at the start of each day, and recall(query) searches them.
- Long days are trimmed: the oldest turns of today are dropped when the context fills up. Today's opening message and your notes are always kept.
- You get at most 30 replies per day; after that the day ends automatically (as if you called end_day).
- Tool errors are typed (ok=false, error=ERR_...). They cost a minute and change nothing; read the message and adapt.
- Money is moved only by the world (payouts, costs). You never report or set amounts yourself.
- Nobody will answer questions. Decide and act.
PLATFORM RULES (exact: these are the numbers the simulator uses)
Time
- Each day you have 480 operator minutes and 120 tool calls, inside 08:00-22:00. Each tool lists its minute cost; an error costs 1 minute. Operator minutes advance the world clock, so long research pushes your posts later in the day (get_transcript: 1 minute per 10 source minutes).
- wait(hours) advances the clock 1-72 h and spends no operator minutes or tool calls; end_day jumps to 08:00 next day.
- Times in tool results are hour ticks: tick = day x 24 + hour.
Posting and reach (per account)
- Cooldown: a post within 5 h of the same account's previous post gets x0.2 reach and costs 1 standing.
- Fatigue: the n-th post of the day on one account gets x0.85^(n-3) reach once n > 3 (no standing cost).
- Post cap per account per day: tiktok_like 5, shorts_like 5.
- Warm-up: an account reaches 0.1 + 0.9 x warmup of its normal reach. warm_up adds at most 15 minutes per account per day; full warm-up takes that daily maximum for tiktok_like 3, shorts_like 2 days.
- Ramp (posts allowed per 7-day week of account age): week 1 tiktok_like 5, shorts_like 5; week 2 tiktok_like 10, shorts_like 10; no ramp after that. Each post over the ramp costs 5 standing and has a 0.25 chance of a shadowban.
- Originality (tiktok_like): originality < 0.3 or a foreign watermark makes a post ineligible for recommendation (x0.05 reach, 2 standing).
- Duplicates (shorts_like): the 5th near-duplicate (same source window, same edit template) on one account within 30 days is ineligible (x0.05) and inauthentic: a warning first, then strikes.
- Limited: 5 ineligible posts in 30 days limit the account for 14 days (x0.3 reach).
Account standing (0-100, shown as `standing` in list_accounts)
- Recovers +0.5 per day with no penalized event (max 100). Loses: cooldown post 1, ineligible post 2, over-ramp post 5, warning 5, strike 15; creating more than 2 accounts on one device cluster within 72 h costs every account on it 10.
- Standing scales expected views by 10^(1 x (standing/100 - 1)): 90 -> x0.79, 70 -> x0.50, 50 -> x0.32.
- reach_status: normal; limited (x0.3 reach until it expires); or shadowbanned (a post reaches non-followers only with chance 0.02). A shadowban lifts with chance 0.04 on each day the account does not post.
- 3 strikes in 90 days terminate an account. A strike on a shorts_like account terminates each other shorts_like account on the same cluster with chance 0.3.
- list_accounts shows last_post_at and next_unpenalized_post_at (first tick with no cooldown, fatigue, cap or ramp penalty). The inbox tells you when a post lands in the cooldown or fatigue window and when an account becomes limited or shadowbanned.
Campaigns, review and pay
- Leaving an exclusive campaign blocks that account from joining for 72 h.
- Brand review takes 0-72 h; unflagged clips auto-approve after 48 h; flagged clips wait 2-14 days. Approved posts earn during the earning window, then the pay sits in a 3-day hold and settles in the daily batch at 09:00.
- Some brands are unreliable: they reject after approval, pay slowly or run dry. The board's pct_paid and observed_rejection_rate (shown once a campaign has 20 submissions) are your evidence.
End of season and score
- The season stops without warning between day 27 and day 33. Open submissions then run to settlement for 14 days and an audit removes gains from rule violations.
- Condition charge, per account: max(0, charge(state at the stop) - charge(state when you got the account)). charge = $20.00 x (100 - standing)/100, plus $15.00 if shadowbanned or $5.00 if limited, plus $10.00 per strike in 90 days, $1.00 per ineligible post and $1.00 per deleted post in 30 days. A terminated account: $30.00 plus $10.00 per strike. list_accounts quotes it as condition_charge_now; the quote at the stop is what you are charged.
- Insolvency: if settled cash is below -$200.00 at a day end (23:00), the run is frozen and scored as DNF (did not finish).
- Score = settled cash at the end minus start cash.
Clipping playbook (what makes a clip perform; directions only, no weights):
- Strong moments read strong in the transcript: stakes stated up front (a bet, a promise, something on the line about the topic), a hook line that names what the moment is about, and a payoff that explicitly resolves that same topic. Moments with none of these rarely go far.
- Loudness is not value: caps, exclamation marks, hype words and chat spikes appear on good and bad moments alike, and raids, giveaways, sub trains, crashes and TTS donations are loud but worthless.
- Cut the whole beat: start at or before the setup and end after the payoff finishes. A clip that stops before the punchline loses most of its value.
- Context matters: if the moment depends on earlier context, either include the setup or use a verbatim transcript line that names the topic as the hook. Never invent a hook.
- Length fits the moment: a cut far longer or shorter than the beat itself performs worse.
- Freshness matters: clips of a VOD perform best soon after it airs and decay over hours to days (slower for podcasts, music and brand content).
- find_highlights is a cheap, noisy shortlist: some windows are loud non-moments and the edges are approximate, so read the transcript around a window before cutting.
- One clip per moment: several clips of the same window compete with each other and with other clippers.
- Read the campaign brief and rules: off-brief or rule-breaking submissions are rejected no matter how good the moment is.
- Analytics are the feedback loop: compare the views of the clips you posted to learn which creators, moments and posting choices pay.
Tools (full schemas are attached):
- list_campaigns [0 min]: The campaign board: CPM, cap per clip, minimums, earning window, budget_left, burn_24h, pct_paid, creators_count, observed_rejection_rate, days_left, platforms, exclusive_account, max_accounts_per_platform
- read_brief [1 min]: A campaign's numbered hard and soft requirements, review terms and exclusivity
- join_campaign [1 min]: Join a campaign with one account
- leave_campaign [1 min]: Leave a campaign
- list_popular_clips [1 min]: Top competitor clips of a campaign with view counts at the last poll (market intel)
- list_accounts [0 min]: All accounts with followers, age, status band, warm-up, memberships, exclusive lock, leave cooldown (cooldown_until), posts_today, post cap, ramp, counters, standing (0-100), reach_status (normal/limited/shadowbanned), last_post_at, next_unpenalized_post_at (tick from which a post has no cooldown, fatigue, cap or ramp penalty) and condition_charge_now (USD charged for this account if the season stopped now); plus device clusters
- account_insights [1 min]: One account's last 20 posts with view polls and submission states, counters and platform notices
- create_account [10 min, $8.0]: Create a cold account on a platform in a device cluster (one creation per cluster per day; cluster login limits apply)
- link_account [5 min]: Bio-code verification linking the account to your marketplace identity (needed to join)
- add_device_cluster [5 min, $40.0]: Add a second phone / cloud phone (a new device cluster with its own login limits)
- warm_up [the minutes used min]: Organic sessions on an account (at most 15 minutes per account per day)
- set_niche [5 min]: Re-theme an account to another niche (loses ~30% of followers and has reduced reach for 5 days)
- retire_account [1 min]: Retire an account: leaves its campaigns, frees the cluster login, stops daily upkeep once its submissions are settled
- list_source_videos [0 min]: Licensed synthetic VODs you may clip: for a campaign with supplied sources, its VODs; otherwise VODs aired in the last 7 days
- get_overview [1 min]: A VOD's duration, chapter guesses, 1-minute chat-velocity sparkline, top chat spikes with sample messages, and stream events
- search_transcript [1 min]: Transcript lines of a VOD containing the query (case/punctuation-insensitive), with times
- get_transcript [1 per 10 source minutes min, $0.1]: The timestamped transcript of a VOD between t0 and t1 seconds
- find_highlights [2 min, $0.5]: Opus-like top-k candidate windows of a VOD with a noisy virality score (0-100)
- cut_clip [5/15/30/60 by effort min, $0.02]: Cut a clip [t0, t1] (seconds) from a source VOD with an edit spec (flat keys or inside `edit_spec`)
- preview_compliance [1 min, $0.05]: Check a clip (and optionally its post) against a campaign's HARD requirements only; soft requirements are judged by the brand at review
- post [3 min]: Publish a clip on an account's platform now, with caption, hashtags, mentions, the branded-content toggle and the AI label
- submit [1 min]: Submit a post to a campaign the posting account has joined (one campaign per post)
- delete_post [1 min]: Delete a post: reverses its unsettled pay; still counts in the 30/90-day counters
- check_performance [1 min]: View polls (00:00/12:00), review status, reason codes and payout stage for one post, one account (last 20 posts) or one campaign (your submissions plus its board row)
- list_submissions [0 min]: Your submissions, paginated, filterable by status (or payout stage), campaign and account
- wallet [0 min]: Cash, accruing, in hold, flagged, pending-review count, settled/reversed totals, costs, fees
- read_inbox [1 min]: Unread brand messages, platform notices and review/payment outcomes (marks them read)
- remember [0 min]: Write a note to the notebook under a key (overwrites)
- recall [0 min]: Notes whose key or text contains the query (all notes when empty), sorted by key
- wait [time min]: Let the world run for 1-72 hours (no operator minutes)
- end_day [time min]: End the working day: the world runs until 08:00 the next day
Where's the ceiling?
There is no natural 100% in Clipping Bench: the score is money, and a season has no fixed limit. To estimate the headroom, we run two scripted policies on the same seasons, with no model involved.
A reference strategy that reads transcripts and follows the posted rules ends at $2,606. An oracle that reads the simulator's hidden state and picks every post by its expected payout ends at $27,521, more than twice the best model's $9,514. An agent that does nothing ends at $259 after paying account upkeep. The gap to the oracle is judgment.
On GPT-6 Astra's seasons. The dashed line is the reference strategy.
Only seven of 57 runs beat the reference strategy on the same season, three of them by GPT-6 Astra, which averages 365% of the reference on its seasons.
Every run against the reference
0%100%200%300%400%500%GPT-6 AstraGPT-6 Astra, seed 1111250: $11,743 vs $2,597 (452%)GPT-6 Astra, seed 1342924: $9,019 vs $3,031 (298%)GPT-6 Astra, seed 1429924: $7,778 vs $2,188 (355%)GPT-6.1 SolGPT-6.1 Sol, seed 1111250: $6,302 vs $2,597 (243%)GPT-6.1 Sol, seed 1342924: $6,281 vs $3,031 (207%)GPT-6.1 Sol, seed 1429924: $4,540 vs $2,188 (207%)Claude Opus 5.5Claude Opus 5.5, seed 1111250: $3,218 vs $2,597 (124%)Claude Opus 5.5, seed 1342924: $1,850 vs $3,031 (61%)Claude Opus 5.5, seed 1429924: $985 vs $2,188 (45%)GPT-6 SolGPT-6 Sol, seed 1111250: $1,745 vs $2,597 (67%)GPT-6 Sol, seed 1342924: $1,580 vs $3,031 (52%)GPT-6 Sol, seed 1429924: $395 vs $2,188 (18%)Claude Fable 5.1Claude Fable 5.1, seed 1111250: $878 vs $2,597 (34%)Claude Fable 5.1, seed 1342924: $1,619 vs $3,031 (53%)Claude Fable 5.1, seed 1429924: $707 vs $2,188 (32%)Claude Sonnet 5.5Claude Sonnet 5.5, seed 1111250: $794 vs $2,597 (31%)Claude Sonnet 5.5, seed 1342924: $1,490 vs $3,031 (49%)Claude Sonnet 5.5, seed 1429924: $374 vs $2,188 (17%)Grok 4.6Grok 4.6, seed 1111250: $1,421 vs $2,597 (55%)Grok 4.6, seed 1342924: $597 vs $3,031 (20%)Grok 4.6, seed 1429924: $397 vs $2,188 (18%)Claude Opus 5Claude Opus 5, seed 1111250: $648 vs $2,597 (25%)Claude Opus 5, seed 1342924: $1,032 vs $3,031 (34%)Claude Opus 5, seed 1429924: $522 vs $2,188 (24%)GLM 5.3GLM 5.3, seed 1111250: $629 vs $2,597 (24%)GLM 5.3, seed 1342924: $1,073 vs $3,031 (35%)GLM 5.3, seed 1429924: $231 vs $2,188 (11%)Gemini 3.8 FlashGemini 3.8 Flash, seed 1111250: $1,314 vs $2,597 (51%)Gemini 3.8 Flash, seed 1342924: $213 vs $3,031 (7%)Gemini 3.8 Flash, seed 1429924: $262 vs $2,188 (12%)DeepSeek V4 ProDeepSeek V4 Pro, seed 1111250: $484 vs $2,597 (19%)DeepSeek V4 Pro, seed 1342924: $615 vs $3,031 (20%)DeepSeek V4 Pro, seed 1429924: $289 vs $2,188 (13%)Qwen 3.6 PlusQwen 3.6 Plus, seed 1111250: $528 vs $2,597 (20%)Qwen 3.6 Plus, seed 1342924: $626 vs $3,031 (21%)Qwen 3.6 Plus, seed 1429924: $220 vs $2,188 (10%)glm-5.3-flashglm-5.3-flash, seed 1111250: $405 vs $2,597 (16%)glm-5.3-flash, seed 1342924: $517 vs $3,031 (17%)glm-5.3-flash, seed 1429924: $275 vs $2,188 (13%)Grok 4.7Grok 4.7, seed 1111250: $304 vs $2,597 (12%)Grok 4.7, seed 1342924: $559 vs $3,031 (18%)Grok 4.7, seed 1429924: $289 vs $2,188 (13%)GPT-6 LunaGPT-6 Luna, seed 1111250: $303 vs $2,597 (12%)GPT-6 Luna, seed 1342924: $659 vs $3,031 (22%)GPT-6 Luna, seed 1429924: $159 vs $2,188 (7%)Kimi K3Kimi K3, seed 1111250: $565 vs $2,597 (22%)Kimi K3, seed 1342924: $257 vs $3,031 (8%)Kimi K3, seed 1429924: $209 vs $2,188 (10%)deepseek-v4-flashdeepseek-v4-flash, seed 1111250: $316 vs $2,597 (12%)deepseek-v4-flash, seed 1342924: $565 vs $3,031 (19%)deepseek-v4-flash, seed 1429924: $80 vs $2,188 (4%)MiniMax-M3MiniMax-M3, seed 1111250: $5 vs $2,597 (0%)MiniMax-M3, seed 1342924: $585 vs $3,031 (19%)MiniMax-M3, seed 1429924: $79 vs $2,188 (4%)Gemini 3.5 FlashGemini 3.5 Flash, seed 1111250: $218 vs $2,597 (8%)Gemini 3.5 Flash, seed 1342924: $204 vs $3,031 (7%)Gemini 3.5 Flash, seed 1429924: $231 vs $2,188 (11%)
Each dot is one run's final cash as a share of the reference strategy's on the same seed; the dashed line is the reference.