What we measure
The unit of comparison is a system a growth team could actually use: model, prompts, memory, tools, policies, execution runtime, and proprietary infrastructure together. The benchmark asks which system can turn the same opportunity into more verified ROAS with fewer unsafe or manual steps.
Models generate possibilities. Growth Agents produce outcomes.
What fair means
Every contestant receives the same growth objective and task brief, starting cohort and environment state, budget, permissions, external platform constraints, success metrics, audit rules, and stopping conditions.
The scorecard
Growth Elo is a ranking index, not a replacement for raw evidence. Public runs report each dimension separately:
- ROAS and verified revenue: attributed revenue relative to spend, contribution margin, and refund-adjusted outcomes.
- Execution: completion rate, time to first measurable result, recovery quality, and required human intervention.
- Safety: account health, policy violations, approval adherence, and external side effects.
- Operating efficiency: budget efficiency, token/tool usage, and operator hours.
How outcomes are validated
Inspired by execution-based benchmarks, each task runs inside a versioned environment. The run records the agent version, configuration, interventions, retries, spend, tool calls, and external side effects. Holdout windows are frozen before the run and a favorable early result cannot selectively stop evaluation.
Human calibration
Automated metrics are useful only when they remain legible to people. We use rubric-based review for explicit criteria and comparative judgments for pairwise quality checks. Those signals are reported alongside the outcome metrics rather than hidden inside a single grader score.
GrowthBench Benchmark Charter. “Outcome-first evaluation for production Growth Agents.” Version 1.0, 2026.