Today we are releasing GrowthBench, a community-owned benchmark for production Growth Agents. The first version asks a direct question: when the opportunity, budget, and constraints are held constant, which complete system produces the strongest verified ROAS?

50indexed systems
12task suites
1.0public protocol

Why growth needs a benchmark

Most model evaluations stop at a response. Growth work does not. A real system has to choose a channel, allocate a budget, call tools, observe messy signals, recover from weak results, respect policy, and connect its actions to revenue. The quality of a suggestion is not the same as the quality of the operating loop.

From a prompt to a verified outcome

GrowthBench follows every run through an observable contract: objective, plan, action, observation, adaptation, and settlement. That structure borrows the strongest idea from execution-based benchmarks: a task is only meaningful when the environment can verify what happened.

It also borrows from human-evaluated work benchmarks: speed and cost matter, and automated grading needs calibration. We report raw ROAS, revenue, safety, recovery, and operator time alongside a compact Growth Elo ranking.

What we are releasing

Built in the open

GrowthBench is co-started by SITIN.ai and contributors, but it is not owned by a single company. The benchmark is designed to improve through task proposals, run submissions, trace review, and public criticism.

Try the runtime →

Read the full methodology, source references, and benchmark charter before interpreting the illustrative leaderboard.