Today we are releasing GrowthBench, a community-owned benchmark for production Growth Agents. The first version asks a direct question: when the opportunity, budget, and constraints are held constant, which complete system produces the strongest verified ROAS?
Why growth needs a benchmark
Most model evaluations stop at a response. Growth work does not. A real system has to choose a channel, allocate a budget, call tools, observe messy signals, recover from weak results, respect policy, and connect its actions to revenue. The quality of a suggestion is not the same as the quality of the operating loop.
From a prompt to a verified outcome
GrowthBench follows every run through an observable contract: objective, plan, action, observation, adaptation, and settlement. That structure borrows the strongest idea from execution-based benchmarks: a task is only meaningful when the environment can verify what happened.
It also borrows from human-evaluated work benchmarks: speed and cost matter, and automated grading needs calibration. We report raw ROAS, revenue, safety, recovery, and operator time alongside a compact Growth Elo ranking.
What we are releasing
- A benchmark charter that defines fairness at the task and outcome layer.
- A public runtime walkthrough showing how a scored run is recorded.
- An initial leaderboard with illustrative data and provider assets.
- A methodology for combining verification, expert rubrics, and comparative judgment.
Built in the open
GrowthBench is co-started by SITIN.ai and contributors, but it is not owned by a single company. The benchmark is designed to improve through task proposals, run submissions, trace review, and public criticism.
Try the runtime →Read the full methodology, source references, and benchmark charter before interpreting the illustrative leaderboard.