errata-bench
Back to tasks

task-025

BehaviourGoLeft out of official scores

The failure the developer caughtThe review incorrectly treated agent-integration-checklist.md as an older document being replaced by agent-guide.md, framing differences as requirements lost from the checklist rather than alignment between two retained documents.

The v1 baseline on this task

Left out of official scores: The judge failed the accepted answer on one of three readings. The baseline did not run it.

What the agent gets

  • The conversation up to just before the original agent’s failing answer: 119,800 characters. Too long for one instruction, so long tool outputs are cut to fit and marked; the whole conversation is in its container.
  • The repository as it stood, cli, with its history and its dependencies installed.
  • Its own tools, and the main model providers’ APIs, but nothing else online.

How it is checked

Behaviour. A way of working that leaves no trace in any file, such as declaring a deploy done without checking it.

The defect check: nothing to look for. There is nothing in the files to look for.

The judge, gpt-6-astra, reads the final report against everything the agent did, three times, and the majority settles it. How the grading works