A run is not a prompt.
It is an operating loop.

GrowthBench evaluates the system that survives the whole path from objective to verified ROAS: planning, tool use, adaptation, recovery, and settlement.

Execution traces / SWE-benchHuman comparison / GDPvalRubrics + preference / JudgementBench

Every run leaves a trace.

Ready to simulate

Objective

Start with a frozen growth brief: target cohort, budget, permissions, deadline, success metrics, and stopping conditions.

Input Growth briefOutput Executable objective

Put the runtime through a run.

GB-TRIAL-001
GrowthBench / simulatorAwaiting objective
Choose a growth objective
Ready00:00
00:00Awaiting objective selection
--:--Agent trace will appear here

Agent vs human baseline

Growth Elohigher is better
52.772.8
ROASverified revenue / spend
2.1x3.86x
Operator hourslower is better
16h6h
Human baselineGrowth Agent

Score what survives contact.

01

Frozen cohort

Starting state and holdout window are locked before the first action.

02

Instrumented effects

Spend, approvals, retries, tool calls, and external side effects are recorded.

03

Verified outcomes

Revenue, refunds, retention, and policy events require a source of truth.

04

Human calibration

Expert rubrics and pairwise comparisons check the automated score.

Outcome-first by design.

Composite Growth Elo is a compact index for ranking. It never replaces the raw measurements that explain why an agent won: ROAS, verified revenue, execution completion, recovery quality, safety, and operator time.

TaskSame opportunityObjective, cohort, budget
AgentComplete systemModel + memory + tools
EvidencePublic traceActions + side effects
ResultVerified outcomeROAS + safety + time
Read the benchmark charter ↗