Abstract
GrowthBench introduces a public evaluation framework for systems that plan, execute, adapt, and settle growth work. Unlike language-model leaderboards, the benchmark treats the complete deployable agent as the unit of comparison. Each contestant receives the same objective, starting cohort, budget, permissions, external constraints, and deadline. The benchmark records the trajectory from objective to plan, tool execution, feedback, recovery, and verified outcome.
Our primary result is outcome-first: ROAS, verified revenue, safety, completion, recovery, and operator time remain visible as raw measurements. Growth Elo provides a compact ranking index, while human rubric and comparative-judgment checks calibrate the quality of the automated score.
GrowthBench measures what survives contact with the real world.
Contributions
- A reproducible benchmark contract for complete Growth Agents rather than isolated models.
- Versioned task suites with frozen cohorts, budgets, permissions, and holdout windows.
- An execution trace format that records actions, retries, approvals, spend, and external effects.
- A public scorecard combining verified ROAS with efficiency, recovery, safety, and human calibration.
Benchmark release
Citation
growthbench2026 / GrowthBench: Which Growth Agent Can Actually Grow the Business? / Version 1.0 / 2026