- Docs
- Run and grade your agent
Run and grade your agent
Test any agent that runs in on the 55 tasks, then grade its reports with your own key. The agent runs on your machine; grading reads the stored records afterwards.
Before you start
- Linux with Docker, Compose v2 and buildx. The official network rule, which lets the agent reach model APIs and nothing else, needs Linux containers and nftables.
- Python 3.12, in two environments: Harbor needs an older OpenAI library than errata-bench’s grader.
- Your own keys: one for your agent’s model, and one for the judge.
- Grading
- About $1.50 an answer. A full run is 51 tasks × 3 attempts = 153 answers: about $230. In the v1 baseline, grading 918 answers cost $1,084.23 at Azure list prices, by the run’s own tally: about $1.18 each.
- Your agent
- Its own model’s cost. With the reference agent, grok-4.6 averaged $0.44 an attempt and DeepSeek-V4-Pro $0.07.
- Time and disk
- The first build of all 55 images takes about 2 hours, and Docker’s build cache grows to about 80 GB. With 4 trials at a time, a full run takes a few hours.
Step 1: Install
Use code v1.0.4 with dataset v1.0.2. This is a code-only release: the tasks, digests and admission are unchanged. Keep your existing v1.0.2 download.
git clone --branch v1.0.4 https://github.com/zanwenfu/errata-bench.git
cd errata-bench
python3.12 -m venv .venv-harbor
.venv-harbor/bin/pip install harbor==0.23.0
.venv-harbor/bin/pip install --no-deps -e .
python3.12 -m venv .venv-grade
.venv-grade/bin/pip install -r requirements-lock.txt
.venv-grade/bin/pip install --no-deps -e .
.venv-grade/bin/pip install huggingface_hubStep 2: Download the tasks
Accept the terms on the dataset page first. The data is for evaluation and research only: no training on it, and no attempt to identify the developers.
source .venv-grade/bin/activate
hf auth login
hf download zanwenfu/errata-bench-v1 --repo-type dataset --revision v1.0.2 --local-dir release/v1.0.2Step 3: Run your agent
Choose an agent, set your key and model, and run from the checkout. Each task is attempted 3 times.
export ANTHROPIC_API_KEY="your-provider-key"
export AGENT_MODEL="anthropic/your-model-id"
.venv-harbor/bin/harbor run -p release/v1.0.2/harbor \
-a claude-code -m "$AGENT_MODEL" \
--ak disable_web_search=true \
-k 3 -n 4 --max-retries 2 -o jobs --job-name my-agentUse the Anthropic model ID your key can access. Provider-side web search must be disabled.
- While the agent works it can reach the main model APIs and nothing else. Claude Code and Codex must turn off their provider-side web search.
- Give each model its own job, with
-nwithin its rate limit.--max-retries 2reruns a trial that failed for the infrastructure. Grade only once every job has ended. - Other Harbor agents work if they write an ATIF trajectory.
Try everything without paying for a model
The stand-in makes a few fixed calls and reports exactly that.
.venv-harbor/bin/harbor run -p release/v1.0.2/harbor \
-a errata_harbor.agents:StandIn -k 1 -n 4 \
-o jobs --job-name stand-inStep 4: Grade
First check which trials can be graded. This reads the records and calls no model.
.venv-grade/bin/python scripts/grade_harbor.py \
release/v1.0.2 jobs/my-agent --out runs/my-agent \
--admission release/v1.0.2/admission/gpt-6-astra --rows-onlyThen grade. gpt-6-astra reads each answer 3 times with both grading methods, and the majority decides. This makes paid calls with your key.
export ERRATA_JUDGE_MODEL="gpt-6-astra"
export OPENAI_API_KEY="your-judge-key"
.venv-grade/bin/python scripts/grade_harbor.py \
release/v1.0.2 jobs/my-agent --out runs/my-agent \
--admission release/v1.0.2/admission/gpt-6-astra- Grading refuses a judge its admission did not check, and a second grading run in the same folder, before paying for anything. Grade a smoke test into a folder of its own.
scripts/harbor_spend.pyprices a run so far, andscripts/harbor-guard.shstops it at a dollar limit you set.
Step 5: Read the results
runs/my-agent/results.json scores each model on its own, with the same measures as the leaderboard.
Reading your results explains every measure, when a result is official, and what has to match before you compare it with the leaderboard.
From the benchmark’s own documentation: docs/running.md · results/v1-baseline