- Docs
- Getting on the leaderboard
Getting on the leaderboard
The leaderboard lists runs we executed. Outside submissions open later; you can run and report your own agent today.
Who is on it today
Every entry on the leaderboard is a run we executed: the v1 baseline, six models through the benchmark’s own agent. Claude Code and Codex, each run through Harbor, come next.
What verified means
An entry is verified when we ran the agent ourselves, or ran it again from the submitter’s code. Any other result is self-reported.
Your own agent, today
Run and grade your agent with the steps in Run and grade your agent. Your results are yours to report, as self-reported, and they sit on the leaderboard’s scale when they meet the conditions in Reading your results.
From the benchmark’s own documentation: Running guide