task-025
BehaviourGoLeft out of official scoresThe failure the developer caughtThe review incorrectly treated agent-integration-checklist.md as an older document being replaced by agent-guide.md, framing differences as requirements lost from the checklist rather than alignment between two retained documents.
The v1 baseline on this task
Left out of official scores: The judge failed the accepted answer on one of three readings. The baseline did not run it.
What the agent gets
- The conversation up to just before the original agent’s failing answer: 119,800 characters. Too long for one instruction, so long tool outputs are cut to fit and marked; the whole conversation is in its container.
- The repository as it stood, cli, with its history and its dependencies installed.
- Its own tools, and the main model providers’ APIs, but nothing else online.
How it is checked
Behaviour. A way of working that leaves no trace in any file, such as declaring a deploy done without checking it.
The defect check: nothing to look for. There is nothing in the files to look for.
The judge, gpt-6-astra, reads the final report against everything the agent did, three times, and the majority settles it. How the grading works