errata-bench
DocsRun and grade your agent
  1. Docs
  2. Run and grade your agent

Run and grade your agent

Test any agent that runs in on the 55 tasks, then grade its reports with your own key. The agent runs on your machine; grading reads the stored records afterwards.

Before you start

  • Linux with Docker, Compose v2 and buildx. The official network rule, which lets the agent reach model APIs and nothing else, needs Linux containers and nftables.
  • Python 3.12, in two environments: Harbor needs an older OpenAI library than errata-bench’s grader.
  • Your own keys: one for your agent’s model, and one for the judge.
Grading
About $1.50 an answer. A full run is 51 tasks × 3 attempts = 153 answers: about $230. In the v1 baseline, grading 918 answers cost $1,084.23 at Azure list prices, by the run’s own tally: about $1.18 each.
Your agent
Its own model’s cost. With the reference agent, grok-4.6 averaged $0.44 an attempt and DeepSeek-V4-Pro $0.07.
Time and disk
The first build of all 55 images takes about 2 hours, and Docker’s build cache grows to about 80 GB. With 4 trials at a time, a full run takes a few hours.

Step 1: Install

Use code v1.0.4 with dataset v1.0.2. This is a code-only release: the tasks, digests and admission are unchanged. Keep your existing v1.0.2 download.

Install the benchmark
git clone --branch v1.0.4 https://github.com/zanwenfu/errata-bench.git
cd errata-bench

python3.12 -m venv .venv-harbor
.venv-harbor/bin/pip install harbor==0.23.0
.venv-harbor/bin/pip install --no-deps -e .

python3.12 -m venv .venv-grade
.venv-grade/bin/pip install -r requirements-lock.txt
.venv-grade/bin/pip install --no-deps -e .
.venv-grade/bin/pip install huggingface_hub

Step 2: Download the tasks

Accept the terms on the dataset page first. The data is for evaluation and research only: no training on it, and no attempt to identify the developers.

Download v1.0.2 tasks
source .venv-grade/bin/activate
hf auth login
hf download zanwenfu/errata-bench-v1 --repo-type dataset --revision v1.0.2 --local-dir release/v1.0.2

Step 3: Run your agent

Choose an agent, set your key and model, and run from the checkout. Each task is attempted 3 times.

Run Claude Code
export ANTHROPIC_API_KEY="your-provider-key"
export AGENT_MODEL="anthropic/your-model-id"

.venv-harbor/bin/harbor run -p release/v1.0.2/harbor \
  -a claude-code -m "$AGENT_MODEL" \
  --ak disable_web_search=true \
  -k 3 -n 4 --max-retries 2 -o jobs --job-name my-agent

Use the Anthropic model ID your key can access. Provider-side web search must be disabled.

  • While the agent works it can reach the main model APIs and nothing else. Claude Code and Codex must turn off their provider-side web search.
  • Give each model its own job, with -n within its rate limit. --max-retries 2 reruns a trial that failed for the infrastructure. Grade only once every job has ended.
  • Other Harbor agents work if they write an ATIF trajectory.
Try everything without paying for a model

The stand-in makes a few fixed calls and reports exactly that.

Run the stand-in
.venv-harbor/bin/harbor run -p release/v1.0.2/harbor \
  -a errata_harbor.agents:StandIn -k 1 -n 4 \
  -o jobs --job-name stand-in

Step 4: Grade

First check which trials can be graded. This reads the records and calls no model.

Check the records
.venv-grade/bin/python scripts/grade_harbor.py \
  release/v1.0.2 jobs/my-agent --out runs/my-agent \
  --admission release/v1.0.2/admission/gpt-6-astra --rows-only

Then grade. gpt-6-astra reads each answer 3 times with both grading methods, and the majority decides. This makes paid calls with your key.

Grade the answers
export ERRATA_JUDGE_MODEL="gpt-6-astra"
export OPENAI_API_KEY="your-judge-key"

.venv-grade/bin/python scripts/grade_harbor.py \
  release/v1.0.2 jobs/my-agent --out runs/my-agent \
  --admission release/v1.0.2/admission/gpt-6-astra

  • Grading refuses a judge its admission did not check, and a second grading run in the same folder, before paying for anything. Grade a smoke test into a folder of its own.
  • scripts/harbor_spend.py prices a run so far, and scripts/harbor-guard.sh stops it at a dollar limit you set.

Step 5: Read the results

runs/my-agent/results.json scores each model on its own, with the same measures as the leaderboard.

Reading your results explains every measure, when a result is official, and what has to match before you compare it with the leaderboard.

From the benchmark’s own documentation: docs/running.md · results/v1-baseline