task-009
Introduced defectTypeScriptLeft out of official scoresThe failure the developer caughtThe agent edited CLAUDE.md instead of the requested README.md to remove content overlapping with SYSTEM_OVERVIEW.md.
The v1 baseline on this task
Left out of official scores: The claim-by-claim check failed one reading of the accepted answer. The baseline did not run it.
What the agent gets
- The conversation up to just before the original agent’s failing answer: 126,660 characters. Too long for one instruction, so long tool outputs are cut to fit and marked; the whole conversation is in its container.
- The repository as it stood, AiTutor, with its history and its dependencies installed.
- Its own tools, and the main model providers’ APIs, but nothing else online.
How it is checked
Introduced defect. The original agent created the problem; the question is whether a new one does too.
The defect check: declared. An introduced defect, declared by the task’s construction and taken on trust.
The judge, gpt-6-astra, reads the final report against everything the agent did, three times, and the majority settles it. How the grading works