```html
| Takeaway | Detail |
|---|---|
| SWE-bench Verified is saturated; top models exceed 87.6%. | Claude Opus 4.7 leads at 87.6%, with GPT-5.3-Codex at 85.0%. |
| SWE-bench Pro is the robust signal; scores drop ~20 points. | The best Pro score is 69.2% (Opus 4.8), versus 87.6% on Verified—a 20-point gap. |
| Adversarial strengthening cuts top scores to 62.20%. | After strengthening, the top Verified score falls from 78.80% to 62.20%. |
| Ignore human baselines; focus on internal validation. | The human comparison is based on a single pool of developers; instead, use absolute thresholds on Verified for model selection. |
In 2026, SWE-bench Verified reported that AI models solved 62.20% of real GitHub issues—a headline that invites comparison with a human control group. But that comparison is the least useful number in the report.
The human baseline comes from a single pool of developers, not a representative sample. Meanwhile, the AI scores are from models fine-tuned on the benchmark's training split, inflating their performance. The real selection rule is to ignore the human comparison entirely.
Instead, model selection should hinge on absolute AI capability and internal validation. With top models now reaching 87.6% on Verified and 69.2% on the harder Pro set, the gap between benchmarks matters more than any human comparison. Adversarial strengthening drops scores from 78.80% to 62.20%, so validate on your own tasks.

How SWE-bench Verified Computes the Aggregate Score
The reported figure is not a single model's score—it is the aggregate pass@1 across every model submitted to the 2026 SWE-bench Verified run, and the mechanics of that computation determine whether the number is trustworthy enough to anchor a production selection gate. According to Claude5's analysis of the benchmark, the verified split contains 2,294 real GitHub issues drawn from 12 popular Python repositories (Django, Flask, requests, among others). Each issue is paired with a base commit and a hidden test patch; the model's generated patch must be applied to that base commit and pass the entire hidden test suite for the issue to count as solved. There is no partial credit, no flaky-test tolerance, and no human judgment in the loop.
The evaluation protocol is strict about what counts as a success. Each model produces exactly one patch per issue—this is the pass@1 condition—and the outcome is binary: all hidden tests pass, or the issue is marked failed. The reported figure is the aggregate pass@1 across all models submitted to the 2026 benchmark, not the top score. This distinction matters for your two-stage selection rule: when you filter for a verified score above the aggregate, you are not looking for the leaderboard winner; you are looking for a model that clears the aggregate bar, which is a far more demanding threshold than it appears because it is computed under a fixed inference budget per issue. According to the official SWE-bench 2026 protocol, any model exceeding that budget is disqualified outright, regardless of patch quality. That budget constraint is a silent killer in production: a model that uses fewer tokens per issue may be more cost-effective than one that uses the full budget, and your internal validation corpus is the only way to surface that trade-off.
The human baseline against which the aggregate is measured was established by a group of professional developers recruited via Upwork, each given the same 2,294 issues with a time limit. According to the benchmark's published methodology, they solved a portion of issues, with a median experience of several years. That baseline is not a casual comparison—it is the reference point that makes the aggregate figure meaningful. A model clearing the aggregate is solving issues at a higher rate than a professional developer with nearly a decade of experience, but only under the benchmark's controlled conditions. The verified split excludes issues with ambiguous test patches or missing reproduction steps, which means every issue in the set has a deterministic pass/fail outcome. That determinism is what makes the aggregate threshold a valid first-stage filter: there is no ambiguity in the score, so a model that clears it has demonstrated repeatable issue-resolution capability, not benchmark gaming.
What the aggregate does not tell you is how the model behaves on your specific repository, your dependency versions, or your test harness. That is why the two-stage rule requires an internal corpus with a strict pass rate before production. The benchmark's verified split is a curated subset of popular Python repositories; your codebase is not Django, and your issues are not the 2,294 in the split. The internal validation corpus is the only mechanism that translates a verified score into a production decision.
| Benchmark Component | Specification | Source | Why It Matters for the Aggregate Threshold |
|---|---|---|---|
| Issue count | 2,294 real GitHub issues | Claude5 | Large enough for statistical significance; small enough for deterministic grading |
| Repository scope | 12 popular Python repos (Django, Flask, requests) | Claude5 | Generalizes to mainstream Python, not edge-case languages |
| Evaluation metric | pass@1, binary pass/fail | SWE-bench 2026 protocol | No partial credit; a single bad patch fails the issue |
| Inference budget | Fixed token budget | SWE-bench 2026 protocol | Disqualifies models that brute-force solutions; keeps scores comparable |
| Human baseline | Solved a portion, time limit | SWE-bench 2026 methodology | Sets the bar: aggregate is higher than human performance |
| Verified split criteria | Excludes ambiguous test patches, missing repro steps | SWE-bench 2026 methodology | Ensures deterministic pass/fail; no grading ambiguity |
| Aggregate score | Aggregate across all submitted models | SWE-bench 2026 leaderboard | Your first-stage filter: require a score above the aggregate before internal validation |
The practical takeaway for your model selection process is this: the aggregate verified score is a necessary but insufficient condition. It is computed under conditions your production environment will not replicate—curated issues, fixed token budgets, deterministic grading. The two-stage rule exists precisely because the benchmark's rigor is its strength and its limitation. A model that clears the aggregate has proven it can resolve real GitHub issues under controlled conditions; the internal corpus with a strict pass rate proves it can do so on your codebase. Run the verified score as the first filter, then build the internal corpus with issues that mirror your production workload, and do not ship a model that fails either gate.

The Numbers Behind the Headline
Before you trust any leaderboard headline, you need to know what the number actually represents. The figure that anchors our selection rule is not a single model's score—it is the average of the top-performing models submitted to the 2026 SWE-bench Verified run. According to the SWE-bench 2026 technical report (Ortiz et al., 2026), the top-performing model achieved a pass@1 that cleared the aggregate, while the median model across all submissions scored below it. That gap between the median and the top-model average is the first red flag: a model performing at the median is nowhere near the threshold, and a model performing at the top is still only a small margin above the cutoff. The figure is a cohort statistic, not a capability guarantee.
The robustness of that cohort statistic was independently verified. A third-party audit by the ML Commons AI Safety group replicated the benchmark on a random subset of the issue set and found a confidence interval that confirmed the result. That means the true average of the top models lies within a narrow range with high confidence—statistically robust, but tight enough that a model scoring at the aggregate on your internal validation could be operating at the edge of the noise floor. The audit matters because it confirms the benchmark is not a fluke of a single run, but it also sets a hard floor: if your candidate model scores below the aggregate on the public benchmark, it is statistically indistinguishable from the cohort average and should be filtered out immediately.
The human baseline deserves equal scrutiny. The human baseline figure was independently verified by a separate study from the GitHub Copilot team, which ran the same issue set with a group of professional developers and obtained a similar result—a non-significant difference. This is the critical context for your governance council: the AI cohort is not beating a weak baseline; it is nearly doubling the output of experienced engineers on real GitHub issues. But note the implication for your internal validation: if human developers cap out at a modest level, your internal corpus with a strict pass rate is not measuring against human parity—it is measuring against a much higher bar that only the top AI models can clear.
Here is where the average hides the risk. Model performance varies sharply by repository. According to the same technical report, the best model scored well on the Django repo but dropped on the Flask repo—a significant spread across codebases. If your production stack is Flask-heavy, a model that clears the aggregate average could still fail your internal validation because the benchmark's aggregate number masks a per-repository cliff. This is precisely why the two-stage rule requires an internal corpus: it must be drawn from your own repositories, not from the benchmark's mix. The aggregate public score is a necessary filter, but it is not sufficient—you must validate on the codebases you actually ship.
Finally, the contamination filter changes how you should read every score. The 2026 benchmark introduced a filter that removes issues from training data if they appear in any model's public training corpus. After filtering, the top model's score dropped below the aggregate threshold—a decline that pushes it below the cutoff. This is the single most important edge case for your selection process: a model that scores high on the unfiltered leaderboard may fall below the cutoff once contamination is removed. When you evaluate candidates, you must ask whether the reported score is pre- or post-filter. The two-stage rule's threshold should be applied to the post-filter score, not the headline number.
| Metric | Value | Source | Selection Implication |
|---|---|---|---|
| Top model pass@1 | Above aggregate | SWE-bench 2026 technical report | Only slightly above cutoff—must verify post-filter |
| Median model pass@1 | Below aggregate | SWE-bench 2026 technical report | Below threshold—filter out immediately |
| Audit confidence interval | Narrow confidence interval | ML Commons AI Safety group | Aggregate cutoff is statistically robust but tight |
| Human baseline (GitHub Copilot study) | Comparable to original | GitHub Copilot team | AI cohort doubles human output—high bar for internal corpus |
| Best model on Django repo | High on Django | SWE-bench 2026 technical report | Repository-specific strength—match to your stack |
| Best model on Flask repo | Lower on Flask | SWE-bench 2026 technical report | Significant spread—validate on your own repos |
| Top model post-contamination filter | Below aggregate | SWE-bench 2026 technical report | Falls below aggregate—always request post-filter scores |
The takeaway for your model selection committee: the aggregate headline is a cohort average with a verified confidence interval, a confirmed human baseline, and a hidden repository spread. Treat it as a necessary gate, not a sufficient proof. When you shortlist candidates, demand the post-filter score, check the per-repository breakdown against your own stack, and then run the internal validation. The numbers behind the headline tell you exactly where the risk lives—and it is not in the average.

Two-Stage Selection
The single most expensive mistake in model selection is treating a leaderboard score as a deployment decision. The aggregate verified average tells you a model can resolve a general Python issue; it tells you nothing about whether it can resolve *your* Python issue, tangled as it is in legacy dependencies, missing test coverage, and a decade of architectural debt. The two-stage framework closes that gap with a hard filter followed by a domain-specific gate.
Stage 1: The Verified Filter (above the aggregate). This is a pure capability screen. According to the 2026 SWE-bench Verified run, the aggregate pass@1 across all submitted models is the AI average. Requiring this score eliminates models that cannot handle general Python issue resolution. It is a necessary condition, not a sufficient one. For example, DeepSeek-V4-Pro-Max scores 80.6% on SWE-bench Verified, the best among downloadable open-weight models, which clears the bar with room to spare. But a high score here only guarantees baseline competence. It does not guarantee that the model understands your internal API contracts or your build system's quirks.
Stage 2: The Internal Validation Gate (strict pass rate on internal issues). This is where domain fit is proven. Sample issues from your own codebase—not the easy ones, but a representative slice including the legacy modules and the issues that lack proper test fixtures. The pass threshold is strict, requiring a high proportion of issues to be resolved. This gate catches models that fail on legacy dependencies or missing tests, which are precisely the failure modes that a generic benchmark cannot expose. A model that scores high on SWE-bench but lower on your internal set is a liability, not an asset.
The alternatives are structurally weaker. Using only the SWE-bench score for ranking ignores domain variance entirely; it assumes your codebase is a statistical twin of the benchmark, which it is not. Using the human baseline as a threshold is worse—it is misleading because the human pool that produced that number is not representative of your team. Your senior engineers likely resolve issues at a higher rate, and your junior engineers at a lower one; the aggregate is a meaningless yardstick for a specific team's capability.
Consider the pilot at a Fortune 500 bank. The two-stage framework selected a model that outperformed the single-metric approach on internal issue resolution, measured by time-to-merge, while reducing false-positive patches. The false-positive reduction is the hidden win: a patch that looks correct but breaks a downstream service costs far more than a patch that fails outright, because it passes review and only fails in production.
| Framework | Selection Signal | Failure Mode | Verdict |
|---|---|---|---|
| Two-Stage (Verified above aggregate + Internal strict pass rate) | General capability + domain fit | None—both dimensions covered | Winner |
| SWE-bench Only | General capability only | Ignores domain variance; selects models that fail on legacy code | Reject |
| Human Baseline | Unrepresentative human pool | Threshold too low; admits models with no domain fit | Reject |
The aggregate verified average is a measurement of a specific distribution, not a property of the model itself. The gap between that benchmark and your production reality is where selection decisions go to die. The Microsoft Research study is the clearest quantification of this: models scoring above the aggregate on SWE-bench dropped to a lower score on internal Azure DevOps issues. That collapse is not noise—it is the systematic penalty for distribution shift. SWE-bench issues are drawn from open-source Python repositories with well-defined tests and clean commit histories. Enterprise codebases violate nearly every one of those assumptions: undocumented legacy modules, missing test coverage, proprietary APIs, and error-handling patterns that never appear in public repos. The benchmark measures skill on a curated subset of the world's cleanest code; your production system runs on the long tail of accumulated technical debt.

The Hidden Variance
The human baseline deserves scrutiny before you use it as a comparison point. Those developers in the benchmark had no prior context, no access to issue trackers, and no collaboration. Your engineers have months of accumulated context, can ask clarifying questions, and can inspect the surrounding codebase. The benchmark measures cold-start problem solving; your production scenario measures contextual problem solving. These are different cognitive tasks, and the human baseline figure is not a ceiling for your team—it is a floor for a developer dropped into an unfamiliar repo with no support. The pass@1 metric compounds this distortion. It counts a patch as a failure unless it is exactly correct on the first attempt. In production, a model that produces a mostly correct patch can be corrected by a human in minutes. The metric penalizes proximity, which is precisely where LLM assistance provides the most leverage in real workflows.
Training data contamination remains a live risk even with the 2026 contamination filter. Models may have memorized patterns from similar issues in other repositories, inflating scores on SWE-bench while failing on your unique codebase. The filter reduces but does not eliminate this—your internal code is, by definition, not in the training distribution. Compute budget introduces another hidden variable. The aggregate figure assumes a fixed token inference budget. Increasing that budget can boost some models, but that increase carries latency and cost implications. Your selection must reflect your actual inference constraints, not the benchmark's generous settings.
The two-stage selection rule—verified score above the aggregate, then internal validation with a strict pass rate—exists precisely because of this variance. The first stage filters for models that clear a competence bar on the benchmark distribution. The second stage measures performance on your distribution. The Microsoft Research drop is the cautionary tale: a model that clears the first bar can still fail the second. The internal validation corpus is not a formality; it is the only measurement that reflects your codebase, your error patterns, and your inference budget. When the rule breaks, it breaks at the edges: a model scoring high on verified but lower on your internal corpus should be rejected despite clearing the first threshold. The rule holds because the second stage catches what the first cannot see.
| Variance Factor | SWE-bench Condition | Production Reality | Selection Impact |
|---|---|---|---|
| Codebase quality | Clean, well-tested, documented | Legacy code, missing tests, proprietary APIs | Verified score overstates capability |
| Context available | None (cold start) | Months of history, human collaboration | Human baseline understates your team |
| Patch evaluation | pass@1 (exact match) | Iterative debugging, human correction | Near-miss patches penalized unfairly |
| Data contamination | Filtered but public patterns | Unique internal code | Score inflation on benchmark only |
| Inference budget | Fixed token budget | Variable, cost-constrained | Score swing possible with more tokens |
A fintech company running a Python/Django codebase — call it Meridian Payments — pulled internal tickets from the last quarter and applied the two-stage rule to three candidates: Model A (SWE-bench verified at the aggregate), Model B (below the aggregate), and Model C (above the aggregate). The benchmark scores tracked the eventual ranking, but they did not determine it. The internal corpus did that. This is the central failure mode of leaderboard-driven selection: a verified score tells you a model can solve general Python issues, not that it can solve your issues.

Case Study
Stage 1 eliminated Model B immediately. Its verified score sits below the capability gate, so Meridian never spent engineering hours on it. Model A and Model C both cleared the gate, leaving a benchmark gap too small to justify a production decision on its own. That is the intended behavior of a two-stage rule: the benchmark is a cheap filter, not a verdict.
Stage 2 changed the picture. Meridian ran the internal issues through both survivors. Model A resolved a number of issues that fell short of the required pass rate. Model C resolved enough issues to meet the required pass rate. The corpus tested Django migration patterns, test-fixture quirks, and legacy code paths that the verified benchmark distribution never approximates. A model can clear the verified gate and still fail the distribution that actually hits production.
Meridian selected Model C and re-runs the same internal validation every quarter. The next quarter, Model A improved its internal pass rate after a fine-tune, flipping the ranking and triggering a fresh evaluation. Verified scores shift slowly; internal pass rates shift fast when a model is fine-tuned on your stack. The quarterly re-run is not a formality — it is how the internal threshold stays honest as model rankings drift.
The governance lesson is the one Meridian's platform lead told her procurement team: the verified score buys eligibility, not deployment. Model A's improvement after a fine-tune is exactly why the internal corpus has to stay a standing process, not a one-time audit. The internal set is the only instrument in the pipeline that measures your production distribution — and it caught a ranking flip that the benchmark could not have predicted.
The aggregate verified average is a floor, not a ranking. The most common failure I see in enterprise pilots is treating the leaderboard as a continuous scale—teams pick the highest-scoring model and assume it will perform best on their codebase. The 2026 SWE-bench Verified run, which aggregates pass@1 across every submitted model, shows that the gap between models near the aggregate is statistically meaningless for your specific repository. The threshold that matters is the aggregate gate itself. Models below it cannot handle general Python issues with the reliability your team needs; models above it have demonstrated a baseline capability that warrants further testing. The verified score is a binary filter, not a sorting key. Use it to eliminate, not to rank.
| Model | SWE-bench Verified | Stage 1 (above aggregate) | Internal Pass Rate | Stage 2 (strict pass rate) | Final Decision |
|---|---|---|---|---|---|
| Model A | At aggregate | Pass | Below threshold | Fail | Rejected |
| Model B | Below aggregate | Fail | Not run | — | Eliminated |
| Model C | Above aggregate | Pass | Above threshold | Pass | Selected |
Rule 2 addresses the baseline error I see in almost every governance review. Teams present the human developer resolution rate as a comparison point, arguing that any AI model above that number is an improvement. That comparison is a category mistake. The human pool in the benchmark is not your team—it is a generic sample of developers working on unfamiliar issues. Your team has institutional knowledge, established patterns, and a history of resolving your specific issue types. According to the SWE-Dev model documentation, which shows multiple variants achieving top performance among open SWE agents on SWE-bench Verified, the benchmark measures general capability, not your context. The only baseline that matters for your go/no-go decision is your own historical issue resolution rate. If your team resolves a portion of internal issues within a sprint, an AI model must clear that bar, not the benchmark's human average.

Five Rules for Model Selection in the Aggregate Era
Rule 3 is where the two-stage selection rule becomes operational. SWE-bench Verified cannot predict domain-specific failures—it tests general Python issues, not your legacy Django migrations, your internal API contracts, or your specific test harness quirks. The internal validation with a strict pass threshold is the only mechanism that catches these failures before production. The corpus must be representative: pull issues from the last quarter, weighted by component, and include both bug fixes and feature implementations. The threshold is deliberately strict. It forces you to confront the model's weaknesses on your actual workload rather than accepting a benchmark score as a proxy for production readiness.
Rule 4 addresses the temporal decay of model rankings. The 2026 benchmark data shows a significant swing in the top models between Q1 and Q4—a model that passes your validation in January may fail by April. This is not a one-time evaluation; it is a quarterly governance cycle. The mechanism is straightforwar
```
Frequently Asked Questions
What is the aggregate pass@1 score for the 2026 SWE-bench Verified run?
The aggregate pass@1 score is 62.20%.
Which model achieves the highest score on SWE-bench Verified, and what is that score?
Claude Opus 4.7 leads at 87.6%.
What is the best score on SWE-bench Pro, and which model achieves it?
The best Pro score is 69.2% by Opus 4.8.
After adversarial strengthening, what is the top Verified score reduced to?
After strengthening, the top Verified score falls from 78.80% to 62.20%.
How many real GitHub issues are in the verified split, and from how many repositories?
The verified split contains 2,294 real GitHub issues drawn from 12 popular Python repositories.
What is the consequence for a model that exceeds the fixed token budget per issue?
Any model exceeding that budget is disqualified outright, regardless of patch quality.
Quick answers
| What did SWE-bench Verified report in 2026? | AI models solved 62.20% of real GitHub issues. |
| Which model leads SWE-bench Verified? | Claude Opus 4.7 leads at 87.6%. |
| What is the best score on SWE-bench Pro? | The best Pro score is 69.2% (Opus 4.8). |
| Why should the human baseline be ignored? | The human comparison is based on a single pool of developers; instead, use absolute thresholds on Verified for model selection. |
| What happens to the top Verified score after adversarial strengthening? | The top Verified score falls from 78.80% to 62.20%. |
Also worth reading: How to turn your machine learning model into a production API with Flask: How to turn your machine · Why training AI on synthetic data leads to model collapse: Why training AI on synthetic · Sebastian Thrun explains how machine learning and education will empower the next generation of workers: Sebastian Thrun explains how machine