Abstract
Language models are increasingly embedded inside systems that plan and execute economically consequential work, but most evaluations continue to measure isolated responses. We introduce GrowthBench, an outcome-first benchmark for production Growth Agents. A Growth Agent is evaluated as a complete deployable system, including its model, prompts, memory, tools, policies, runtime, and proprietary infrastructure.
Each contestant receives the same growth objective, frozen starting cohort, budget, permissions, platform constraints, success metrics, audit rules, and deadline. Runs are instrumented from objective through planning, tool execution, feedback observation, adaptation, recovery, and settlement. The primary outcome is verified ROAS, reported alongside revenue, execution completion, safety, operator time, and public trajectory evidence.
GrowthBench combines execution-based validation with human calibration. Explicit rubrics support auditable criteria, while pairwise comparative judgments test whether the score agrees with expert preferences. Growth Elo provides a compact ranking index but never replaces the raw cohort-level measurements.
Subjects
Artificial Intelligence; Multi-Agent Systems; Human-Computer Interaction; Machine Learning
Release materials
Submission history
v1: Initial public methodology, illustrative leaderboard, and runtime simulator. A formal arXiv identifier will replace the placeholder in the page title when the paper is submitted.
GrowthBench Contributors. GrowthBench: Which Growth Agent Can Actually Grow the Business? Preprint, 2026.