- Docs
- Versions
Versions
The current code is v1.0.4, with dataset v1.0.2. Grade a trial with the dataset folder of the version it ran.
Safeguards before the first official run
Code only: keep the dataset you have. A model’s results are official only when complete, with no answer missing. An answer is settled by the readings whose quotes are in it, when they are most of the readings and agree. A run with other limits or settings is not official. The reference agent keeps credentials out of its commands and stops itself inside Harbor’s time limit, with its record kept. Grading refuses a judge that has not passed admission, and a spend guard can stop a run.
More precise grading records
Code only: keep the dataset you have. The claim-by-claim check confirms where a cut it cites actually occurred. Results name the dataset release and the model that served each grading request.
Grading reads what admission read
Grading now reads each conversation uncut, as admission did. Results record where they came from and which models served them. Windows paths and defect-check labels were corrected.
Model APIs widened to the main providers
Agents using the main model providers can now run officially. Task instructions, images and the verifier stayed the same; task digests changed with the network settings.
The first public release
Runnable Harbor tasks, frozen grading data and the official judge’s admission for each task, released on Hugging Face under the dataset’s access terms.
From the benchmark’s own documentation: docs/running.md