SWE-sweep
How many bugs can LMs find & fix in large codebases?
Given a real repository, an agent must discover & repair as many bugs as they can. Agents are not given any hint about the type of bug or its location.
| Rank | Model | Agent | Bugs resolvedScore | Total USDUSD | USD / repo | Turns / repo | Tokens / repo | |
|---|---|---|---|---|---|---|---|---|
| 1 | Sol 5.6 (xhigh)OpenAI | mini-SWE-agent | 4.7% | $7,230 | $72.30 | 233 | 104.8k | |
| 2 | Luna 5.6 (xhigh)OpenAI | mini-SWE-agent | 2.5% | $224 | $2.24 | 204 | 61.2k | |
| 3 | Terra 5.6 (xhigh)OpenAI | mini-SWE-agent | 1.5% | $357 | $3.57 | 76 | 46.6k | |
| 4 | Luna 5.6 (high)OpenAI | mini-SWE-agent | 1.4% | $28 | $0.28 | 75 | 20.3k | |
| 5 | Opus 5 (xhigh)Anthropic | mini-SWE-agent | 1.3% | $5,363 | $53.63 | 323 | 192.7k | |
| 6 | ![]() |
Kimi K3Moonshot AI | mini-SWE-agent | 0.6% | $2,451 | $24.51 | 337 | 152.6k |
| 7 | Luna 5.6OpenAI | mini-SWE-agent | 0.5% | $4 | $0.04 | 22 | 4.5k | |
| 8 | GPT-5.4 Mini (high)OpenAI | mini-SWE-agent | 0.5% | $122 | $1.22 | 75 | 40.5k | |
| 9 | GPT-5.4 MiniOpenAI | mini-SWE-agent | 0.2% | $5 | $0.05 | 13 | 2.1k | |
| 10 | Gemini 3.5 Flash LiteGoogle | mini-SWE-agent | 0.1% | $6 | $0.06 | 34 | 5.2k |
Per-repository resource use includes non-deprecated retries and is averaged over evaluated repositories.
About
Most existing software engineering benchmarks evaluate coding agents on concrete, well-specified tasks, commonly by providing a codebase together with a user-reported issue to resolve. However, as users delegate increasingly broad outcomes to coding agents, the natural next step is for agents to determine not only how to perform useful work, but also what useful work needs to be done.
An agent entrusted with a repository should be able to decide what is broken, which problems matter, and how to solve them before they are reported. We introduce SWE-sweep, a benchmark for this open-ended setting.
Given a codebase containing many concurrent bugs and no information about their nature or location, an agent must autonomously discover and fix as many bugs as possible.
SWE-sweep is constructed from open-source repositories. For each repository, we collect issue-pull request pairs, then identify a single commit where the maximum number of bugs are present at the same time.
Each repair is evaluated against hidden tests from the corresponding pull requests, along with the existing test suite to check for regressions.
Success requires agents to explore and understand a large codebase over long horizon work, repair bugs without introducing regressions, and manage interactions among fixes that are not independent.
How are tasks constructed?
We collect real issue–pull request pairs, identify a commit where many of those bugs coexist, and retain bugs whose fixes and tests can be reproduced at that repository state. Besides many quality filters shared with other benchmarks, we apply extensive filtering to evaluate only bugs that can be discovered from reading the repository alone.
What does an agent receive?
A repository at a fixed base commit and a broad instruction to find and fix as many bugs as possible. It receives no issue descriptions, filenames, line ranges, or other bug-specific hints.
You can find the full prompt here.
How is SWE-sweep evaluated?
We score every task against a reference set of previously identified bugs (see Construction, counting how many the agent successfully repairs.
For every task, we run the agent's submitted codebase against two sets of tests. First, we restore the repository's original test suite and run it to verify no existing behavior was broken. Second, for each bug, we run a set of hidden tests; at least one of these tests fails on the unmodified codebase, and passes once the bug is fixed (fail-to-pass). A bug is considered resolved if all its hidden tests pass and the original suite still passes.
The benchmark score is the fraction of all bugs across all repositories that have been resolved. What about any other changes that the agent makes?
What bugs are in the benchmark? How do you guarantee the task is feasible?
We filter bugs (represented by a test patch and a fix patch) to ensure the task is feasible. The criteria are:
- The test patch and fix patch independently apply to the base commit.
- All target tests pass after the fix patch has been applied to the base commit, but at least one target test fails on the base commit (F2P tests). There might be additional tests that pass before and after the fix patch has been applied (P2P tests).
- The F2P tests reveal a single discoverable bug in the base commit.
- The target tests are not overly specific; any reasonable fix to the discoverable bug will pass the target tests.
- The target tests do not contradict the base commit tests.
- Target tests of different bug instances do not contradict each other.
Appendix A.2 in the paper discusses feasibility in detail.
What makes a bug discoverable?
The expected behavior must be inferable from the repository itself, for example through documentation, types, existing tests, callers, invariants, standards, or an unambiguously undesirable failure such as a crash or data loss. The latter category is used extremely conservatively and all but 2 bugs in the benchmark have concrete repository contracts that describe the expected behavior. You can find some examples about what we mean with repository contracts here. We have spent a lot of time validating this aspect of the benchmark and you can find more details in the appendix of our paper.
What about any other changes that the agent makes?
Any history-derived benchmark necessarily under-counts the bugs present in a repository. By restricting PandoraBench to defects confirmed by an upstream fix, we ensure that every bug in the benchmark is backed by strong evidence that the observed behavior was considered erroneous by the repository maintainers. We therefore only score the agent's changes on the bugs that are confirmed by an upstream fix and supported by executable regression tests, as well as the other quality filters. However, if an agent causes a regression in the original test suite (the agent is explicitly told to avoid this), it will be scores as 0%. This means that the agent's changes that are not scored by the set of bugs are still likely to be non-destructive and compatible with the repository's existing behavior. See A.3 and A.4 in the paper for more discussion.
Does more inference-time compute help?
Repeated attempts recover additional bugs, but the gains diminish. Later work within one run can also undo earlier repairs, so simply extending a trajectory does not guarantee improvement. See Fig. 7 in the paper.
What about Astra, Fable, 5.5, ...?
We're working on evaluating more models! The current selection was finalized for our ICLR submission. We're also looking into even higher reasoning modes, but this might push over $10k for a single run. We also want to have more open weights models on the leaderboard.
Why mini-swe-agent? Could other scaffolds/multiagents achieve higher performance?
Our paper has an ablation with Claude Code and Codex. Neither seems to significantly outperform mini-swe-agent (to the contrary, mini-swe-agent is even quite a bit better than Codex). This follows many other benchmarks, where mini-swe-agent has been extremely competitive. However, we absolutely hope to kick off more research into the role of agent scaffolds and will open for submissions soon.
How do I submit to the leaderboard?
Public submissions are coming soon.
Citation
@misc{lieret2026swesweep,
title = {{SWE-sweep}: Can Agents Autonomously Find and Fix Bugs?},
author = {Kilian Lieret and Jeffrey Jian Ma and Rahul Kindi and
Yuxiang Wei and Jeremy Ma and Sten Sootla and
Parth Thakkar and Chao Beyond Zhou and Pengcheng Yin and
Rui Hou and Ofir Press and John Yang},
year = {2026},
note = {Preprint}
}
