errata-bench
DocsTroubleshooting
  1. Docs
  2. Troubleshooting

Troubleshooting

What goes wrong most often, and what to do about it.

Hugging Face returns 401 or 403

Accept the dataset’s terms in your browser, then sign in with a token from the same account. Public does not mean ungated.

Harbor fails its network check

The official network rule needs Linux containers, nftables, Docker Compose v2 and buildx. Harbor checks this before a run; NetworkCheck shows from inside a container which hosts are reachable.

A trial is not gradable

Read the reason --rows-only prints. A provider error, a failed build or a missing trajectory is not an answer; run those trials again with harbor jobs resume, naming each error type.

Why only 51 tasks count

All 55 tasks are downloaded and run. A task counts only after the grading model reads its known right and wrong answers and three test answers correctly. The judge’s readings failed on two tasks; the claim-by-claim check’s readings failed on two more. The remaining 51 tasks contribute to official scores.

Can I use another judge?

Yes. Admit it first with scripts/admit_judge.py (about $4 a task) and grade with that admission. Its results are that judge’s, not official.

From the benchmark’s own documentation: docs/running.md