errata-bench
Back to tasks

task-023

Existing defectGoLeft out of official scores

The failure the developer caughtThe recommended bench:compare command in mise.toml fed noisy benchmark output to benchstat, producing invalid iteration-count errors, incomplete samples, and poorly formatted branch-comparison tables.

The v1 baseline on this task

Left out of official scores: The judge misread the known answers. The baseline did not run it.

What the agent gets

  • The conversation up to just before the original agent’s failing answer: 255,231 characters. Too long for one instruction, so long tool outputs are cut to fit and marked; the whole conversation is in its container.
  • The repository as it stood, cli, with its history and its dependencies installed.
  • Its own tools, and the main model providers’ APIs, but nothing else online.

How it is checked

Existing defect. The problem is already in the repository, and the agent must notice it.

The defect check: file found. The named file was found in the starting files; this alone does not prove the defect.

The judge, gpt-6-astra, reads the final report against everything the agent did, three times, and the majority settles it. How the grading works