What we measure

The unit of comparison is a system a growth team could actually use: model, prompts, memory, tools, policies, execution runtime, and proprietary infrastructure together. The benchmark asks which system can turn the same opportunity into more verified ROAS with fewer unsafe or manual steps.

Models generate possibilities. Growth Agents produce outcomes.

What fair means

Every contestant receives the same growth objective and task brief, starting cohort and environment state, budget, permissions, external platform constraints, success metrics, audit rules, and stopping conditions.

Same taskFrozen brief, cohort, deadline, and budget.
Complete agentInternal model and runtime remain part of the product.
Verified resultSpend, revenue, refunds, and safety events need a source of truth.

The scorecard

Growth Elo is a ranking index, not a replacement for raw evidence. Public runs report each dimension separately:

How outcomes are validated

Inspired by execution-based benchmarks, each task runs inside a versioned environment. The run records the agent version, configuration, interventions, retries, spend, tool calls, and external side effects. Holdout windows are frozen before the run and a favorable early result cannot selectively stop evaluation.

Human calibration

Automated metrics are useful only when they remain legible to people. We use rubric-based review for explicit criteria and comparative judgments for pairwise quality checks. Those signals are reported alongside the outcome metrics rather than hidden inside a single grader score.

SWE-benchExecution-based validation and reproducible environments.Read reference ↗
GDPvalHuman comparison, speed, cost, and expert calibration.Read reference ↗
JudgementBenchRubric and comparative-judgment methodology.Read reference ↗

GrowthBench Benchmark Charter. “Outcome-first evaluation for production Growth Agents.” Version 1.0, 2026.