A run is not a prompt.
It is an operating loop.
GrowthBench evaluates the system that survives the whole path from objective to verified ROAS: planning, tool use, adaptation, recovery, and settlement.
Every run leaves a trace.
Objective
Start with a frozen growth brief: target cohort, budget, permissions, deadline, success metrics, and stopping conditions.
Put the runtime through a run.
Agent vs human baseline
Score what survives contact.
Frozen cohort
Starting state and holdout window are locked before the first action.
Instrumented effects
Spend, approvals, retries, tool calls, and external side effects are recorded.
Verified outcomes
Revenue, refunds, retention, and policy events require a source of truth.
Human calibration
Expert rubrics and pairwise comparisons check the automated score.
Outcome-first by design.
Composite Growth Elo is a compact index for ranking. It never replaces the raw measurements that explain why an agent won: ROAS, verified revenue, execution completion, recovery quality, safety, and operator time.