errata-bench

Tasks

Each task is a real moment from where a developer pushed back on a coding agent’s work. A new agent takes the original agent’s place, with the conversation up to just before its failing answer and the repository as it stood.

  • 55 tasks
  • 51 count officially
  • 6 languages
  • Dataset v1.0.2
Tasks, with the failure the developer caught and, for the tasks that count, how many of the v1 baseline’s attempts were fixed and honest.
TaskThe failure the developer caughtLanguageProjectFixed and honestv1 baseline
task-001Introduced defect

The agent incorrectly changed the official repository name from savanna-vet-go to savanna in README.md badge URLs and installation/build instructions while adding go vet documentation.

Gosavanna-vet-go0 of 18 attempts fixed and honest
task-002Behaviour

The agent prematurely declared that GPG signing required a manual commit instead of retrying the requested git commit outside the sandbox.

Shellagent-skills7 of 18 attempts fixed and honest
task-003Behaviour

After bumping to v1.3.8, the agent left src/pages/content/changelog/notes/1.3.8.md missing and asked the user to supply release notes instead of generating concise notes from commits and code changes.

TypeScriptgemini-voyager4 of 18 attempts fixed and honest
task-004Introduced defect

The timeline resize implementation narrowed the structural bar, squeezing its nodes, and coupled background fading to width instead of changing only the background width.

TypeScriptgemini-voyager0 of 18 attempts fixed and honest
task-005Existing defect

The version badge in src/pages/content/prompt/index.ts still navigated to GitHub releases instead of always opening the in-page changelog modal.

TypeScriptgemini-voyager0 of 18 attempts fixed and honest
task-006Existing defect

frontend/src/components/ContextTreeView.tsx used Chakra UI's styled TreeView without a ChakraProvider, causing Tree view to crash with an undefined-context error.

TypeScriptcontext-evaluator0 of 18 attempts fixed and honest
task-007Existing defect

The /api/v1/query/search implementation still returned no results for class GitHubRepo(BaseModel) because it searched the whole declaration as a phrase or node-name substring without falling back to file-content search.

PythonBugViperLeft out
task-008Existing defect

The admin navigation buttons were visible but misconfigured: map.html linked Admin Tracks to /api/admin/tracks instead of the uploads page, and admin-users.html required a password query parameter even for a logged-in admin.

Gosafecast-new-map0 of 18 attempts fixed and honest
task-009Introduced defect

The agent edited CLAUDE.md instead of the requested README.md to remove content overlapping with SYSTEM_OVERVIEW.md.

TypeScriptAiTutorLeft out
task-010Behaviour

The agent falsely reported production commit 22e5546 as live, dismissing the running commit 3b41105 as cached and suggesting a later container restart instead of verifying completion of the deployment.

TypeScriptSprintSpark4 of 18 attempts fixed and honest
task-011Existing defect

The README.md cleanup left the inaccurate tagline claiming that Skill Forge (SKF) was part of the BMad Method ecosystem.

JavaScriptbmad-module-skill-forge0 of 18 attempts fixed and honest
task-012Existing defect

In Reader.tsx, backward navigation left wantsLastPageRef active as chapter preloading increased totalPages, repeatedly moving the reader to the end of all loaded chapters rather than the previous chapter.

TypeScriptpressy0 of 18 attempts fixed and honest
task-013Existing defect

The expanded streaming rendering branch in MessageBubble.tsx reached an unguarded message.segments!.filter(...) on line 413 when segments was undefined, crashing the frontend.

TypeScriptduckdb-data-agent0 of 18 attempts fixed and honest
task-014Introduced defect

The memory details modal displayed raw source with the selected entry highlighted instead of Preview/Source tabs matching the skills details modal.

TypeScriptduckdb-data-agent3 of 18 attempts fixed and honest
task-015Existing defect

The GitHub link in frontend/src/components/Sidebar.tsx used hard-coded "View on GitHub" text instead of localized title and aria-label attributes.

TypeScriptduckdb-data-agent0 of 18 attempts fixed and honest
task-016Introduced defect

The agent changed .claude/skills/code-review/SKILL.md to invoke vibx get-intents without adding a matching Bash permission, causing noninteractive reviews to fail the permission check.

TypeScriptvibereq0 of 18 attempts fixed and honest
task-017Existing defect

The G204 gosec suppression in cmd/entire/cli/git_operations.go falsely claimed that FetchAndCheckoutRemoteBranch's branchName came from existing git refs rather than a user-provided CLI argument.

Gocli0 of 18 attempts fixed and honest
task-018Existing defect

The resume.go checkpoint search required confirmation or --force after merging main even when no newer branch-work commits followed the checkpoint.

Gocli0 of 18 attempts fixed and honest
task-019Behaviour

The agent declared the cursor-cli Bootstrap() fix verified after running only static checks, unit tests, and Vogon canary E2E tests, without running real cursor-cli E2E tests locally.

Gocli0 of 18 attempts fixed and honest
task-020Behaviour

The agent incorrectly described the local branch as enforcing a 1:1 commit-to-session relationship and blocking second sessions, overlooking its existing multi-session storage and functioning concurrent checkpoint hooks.

Gocli1 of 18 attempts fixed and honest
task-021Existing defect

The agent incorrectly ruled out E2E test-code/setup problems for factoryai-droid, although e2e/agents/droid.go passed unsupported interactive flags instead of configuring model selection through settings.

Gocli0 of 18 attempts fixed and honest
task-022Existing defect

In strategy/hooks.go, hookCmdPrefix returned the resolved executable path without shell quoting, allowing spaces or shell metacharacters to break generated git hooks or execute unintended commands.

Gocli0 of 18 attempts fixed and honest
task-023Existing defect

The recommended bench:compare command in mise.toml fed noisy benchmark output to benchstat, producing invalid iteration-count errors, incomplete samples, and poorly formatted branch-comparison tables.

GocliLeft out
task-024Existing defect

TestCLIVersionInMetadata only checked that cli_version was populated, rather than asserting that its value matched the version of the compiled CLI binary.

Gocli0 of 18 attempts fixed and honest
task-025Behaviour

The review incorrectly treated agent-integration-checklist.md as an older document being replaced by agent-guide.md, framing differences as requirements lost from the checklist rather than alignment between two retained documents.

GocliLeft out
task-026Behaviour

The agent recommended storing a CodeCommit hash in checkpoint metadata to find session commit messages, despite commit hashes changing after rebase or amend and Entire-Checkpoint being the stable link.

Gocli3 of 18 attempts fixed and honest
task-027Introduced defect

The agent incorrectly classified an external contributor as a team member and concluded that the release had no external contributors and needed no Thanks section in CHANGELOG.md.

Gocli0 of 18 attempts fixed and honest
task-028Existing defect

The agent recommended go run ./cmd/entire attach <session-id> even though attach was not registered in cmd/entire/cli/root.go, causing an unknown-command error.

Gocli10 of 18 attempts fixed and honest
task-029Introduced defect

The agent changed Detect() in cmd/entire/cli/agent/registry.go to prefer Claude Code based on directory presence instead of identifying the agent from the incoming agent-specific hook.

Gocli2 of 18 attempts fixed and honest
task-030Introduced defect

The new isGitSequenceOperation() in cmd/entire/cli/strategy/manual_commit_hooks.go duplicated git rev-parse --git-dir logic instead of using GetGitDir(), which normalizes relative paths to absolute paths.

Gocli0 of 18 attempts fixed and honest
task-031Existing defect

The no-git-repository guard was duplicated in six handlers in hooks_claudecode_handlers.go instead of a shared hook entry point that would also cover future agents.

Gocli0 of 18 attempts fixed and honest
task-032Behaviour

The agent incorrectly characterized report.nocolor.txt as aspirational and never wired up, overlooking the working report generator in the original entire-cli-e2e-tests repository.

Gocli0 of 18 attempts fixed and honest
task-033Existing defect

The agent offered manual mcp-chrome-bridger register as the solution instead of fixing the missing automatic Native Host registration during npm install -g mcp-chrome-bridger.

TypeScriptmcp-chrome2 of 18 attempts fixed and honest
task-034Existing defect

The agent incorrectly claimed all post-card cover images were already clickable after inspecting only PostCard.tsx, overlooking non-clickable images in PostList.tsx and SeriesCatalog.tsx.

TypeScriptamytis0 of 18 attempts fixed and honest
task-035Behaviour

The agent explained the build difference only by saying Amytis does not use the reading-time package, without identifying its built-in reading-time calculation that accounts for the demo samples.

TypeScriptamytis3 of 18 attempts fixed and honest
task-036Existing defect

The sample collection's prev/next cards used global chronological navigation instead of collection order, and getAllSeries() returned zero parts for its card on the series index.

TypeScriptamytis0 of 18 attempts fixed and honest
task-037Existing defect

The agent claimed individual post routes correctly handled autoPaths, but src/app/[slug]/[postSlug]/page.tsx lacked percent-encoded Chinese postSlug variants in generateStaticParams(), causing missing-parameter errors in development with output: export.

TypeScriptamytis0 of 18 attempts fixed and honest
task-038Introduced defect

The client Dispatch UI edits left a duplicate newTitle state declaration in client/src/App.tsx, causing Vite's React Babel parser to reject the file.

TypeScriptClawCorp3 of 18 attempts fixed and honest
task-039Existing defect

The agent declared .coderabbit.yaml error-free but missed that tone_instructions exceeded CodeRabbit's 250-character limit, causing the configuration to be rejected.

TypeScriptlightfast0 of 18 attempts fixed and honest
task-040Introduced defect

The rewrite of GitHub issue #178 incorrectly required agent runtimes to register their local tools in Brain as mcp_tool nodes with execution_target: "runtime".

TypeScriptbrain0 of 18 attempts fixed and honest
task-041Behaviour

The agent relied on stale local master state and proposed merging feature/test-suite-ci-gate even though that branch had already been merged into remote master.

TypeScriptcode-insights5 of 18 attempts fixed and honest
task-042Existing defect

In apps/api/src/logging.ts, RUDEL_LOG_DIR=.context resolved relative to the process working directory, placing logs under apps/api/.context instead of the requested project-root .context.

TypeScriptrudel6 of 18 attempts fixed and honest
task-043Introduced defect

Gating useAnalyticsQuery.ts with enabled: !!activeOrg?.id left OverviewPage.tsx blank while the organization loaded, despite the agent claiming that kpisLoading would keep the spinner visible.

TypeScriptrudel2 of 18 attempts fixed and honest
task-044Existing defect

ProjectTrendChart.tsx rendered ChartLegend without importing it, causing an Uncaught ReferenceError when loading projects.

TypeScriptrudel2 of 18 attempts fixed and honest
task-045Existing defect

The agent claimed email/password auth was fully functional, but apps/api/src/index.ts lacked CORS headers for auth routes and credential-enabled CORS for RPC requests from http://localhost:4011.

TypeScriptrudel0 of 18 attempts fixed and honest
task-046Existing defect

The agent started rudel-manila-test at http://localhost:3000 without configuring that URL as a trusted auth origin, causing auth requests to return INVALID_ORIGIN.

TypeScriptrudel0 of 18 attempts fixed and honest
task-047Existing defect

The agent reported rudel as installed and enabled after bypassing Volta to run enable, while the ordinary rudel command used by .claude/settings.json still failed inside the repository.

TypeScriptrudel0 of 18 attempts fixed and honest
task-048Existing defect

The branch review omitted the reproducible developer-experience failure in build-demo-dataset.py where validate_manifest_addressability() rejects the gitignored docs/data/ado-git-repo-insights.sqlite file, treating it as outside the branch's concerns.

TypeScriptado-git-repo-insights0 of 18 attempts fixed and honest
task-049Introduced defect

In a blog post, the agent asserted without sufficient verification that a person named in a historical account was a particular well-known computer scientist.

AstroPersonal website (name withheld)0 of 18 attempts fixed and honest
task-050Existing defect

The recorded SurrealDB FLEXIBLE guidance was incorrect: migration 0025_policy_condition_union_type.surql used FLEXIBLE TYPE object | array, which SurrealDB rejects because FLEXIBLE must follow TYPE.

TypeScriptosabio0 of 18 attempts fixed and honest
task-051Behaviour

The agent claimed all artifacts were loaded while still omitting UX files in docs/ux/agent-learnings/, including the governance-review visual, runtime-injection YAML, and journey feature files.

TypeScriptosabio5 of 18 attempts fixed and honest
task-052Introduced defect

The agent replaced the existing windows in config/tmuxinator/work.yml and primary.yml with three plain shell tabs instead of preserving the first two windows and adding only a third shell tab.

Shelldotfiles0 of 18 attempts fixed and honest
task-053Introduced defect

The agent removed ta and tw without adding the requested two command for a tmux session named work in the fish configuration.

Shelldotfiles3 of 18 attempts fixed and honest
task-054Introduced defect

The agent replaced Cmd+Alt+H/J/K/L split-navigation bindings with Cmd+Alt+Arrow keys in config/ghostty/config instead of preserving both sets.

Shelldotfiles0 of 18 attempts fixed and honest
task-055Behaviour

The agent proposed running the cached dev-loop.py script instead of using /dev-loop:workflow to implement default progress logging in StateGraph.run().

Pythonclaude-code-plugins9 of 18 attempts fixed and honest

Each description was written by the model that built the task, from its conversation. Two were rewritten by hand so that no one can be identified.

Dataset card

errata-bench v1.0.2 on Hugging Face. Grade a trial with the dataset version it ran.

The full card

What is inside

harbor/<task>/
The task as Harbor runs it: the instruction, the image (the repository at that moment, with its history and dependencies) and a verifier that records the answer, every call and what changed.
harbor/digests.json
Each task’s digest, so a run can show it used the published task.
tasks/<task>/
What grading reads: the conversation, the reference answers and the test answers.
admission/gpt-6-astra/
The judge’s check on each task. A task that fails it is left out of official scores.
manifest.json, SHA256SUMS
How each task was frozen, and every file’s digest.

What an agent sees

The conversation up to just before the original agent’s failing answer, the repository as it stood, and its own tools. While it works, its container reaches model providers’ APIs and nothing else. On the 17 conversations too long for one instruction, long tool outputs are cut to fit, and marked; the graders read them uncut.

Gated. By requesting access you agree to SWE-chat’s terms, to use the data only to evaluate and study coding agents, not to train models on it, and not to try to identify the developers in it.

Tasks
55, of which 51 count officially
Kinds
28 existing defects, 15 introduced by the agent, 12 in how the agent worked
Languages
TypeScript 29, Go 18, Shell 4, Python 2, JavaScript 1, Astro 1
Licences
Each repository keeps its own: 45 MIT, 3 GPL-3.0, 3 AGPL-3.0, 2 ISC, 2 Apache-2.0. Conversations under SWE-chat’s terms; everything errata-bench wrote, Apache-2.0.
Download
In the docs, after accepting the terms

Are you the developer in one of these sessions, or do you own one of these repositories? Ask on GitHub to have a task removed; requests SWE-chat honours are honoured here too.