errata-bench
Back to tasks

task-009

Introduced defectTypeScriptLeft out of official scores

The failure the developer caughtThe agent edited CLAUDE.md instead of the requested README.md to remove content overlapping with SYSTEM_OVERVIEW.md.

The v1 baseline on this task

Left out of official scores: The claim-by-claim check failed one reading of the accepted answer. The baseline did not run it.

What the agent gets

  • The conversation up to just before the original agent’s failing answer: 126,660 characters. Too long for one instruction, so long tool outputs are cut to fit and marked; the whole conversation is in its container.
  • The repository as it stood, AiTutor, with its history and its dependencies installed.
  • Its own tools, and the main model providers’ APIs, but nothing else online.

How it is checked

Introduced defect. The original agent created the problem; the question is whether a new one does too.

The defect check: declared. An introduced defect, declared by the task’s construction and taken on trust.

The judge, gpt-6-astra, reads the final report against everything the agent did, three times, and the majority settles it. How the grading works