ModelRefs / SWE-Bench Verified Leaderboard — AI Model Scores

SWE-Bench Verified Leaderboard — AI Model Scores

Resolve real GitHub issues end-to-end with passing test suite. Current leaders, methodology, and citation sources for SWE-Bench Verified.

Overview

Resolve real GitHub issues end-to-end with passing test suite.

How it is measured: Human-verified subset; resolved-rate on full repository context.

What this benchmark measures

  • repository-level software issue resolution
  • patch validation

Relevant to:

  • coding-agent evaluation
  • repository maintenance evaluation
  • evaluation-pipeline and observability regression-case design

Failure modes it exercises:

  • incorrect patches
  • test regressions
  • failure to resolve a documented issue

Method and its limits

Resolved-rate on a human-verified subset of real repository issues using repository context and test validation.

  • Results depend on the agent scaffold, tools, execution environment, and inference budget.
  • Passing benchmark tests does not establish security, maintainability, or review quality.

Dataset

Dataset
SWE-bench Verified
Type
real GitHub issues with repository test suites
Freshness
aging

The record identifies the 500-instance SWE-bench Verified release. Repository, language, and issue-type coverage should be checked at the source; the release date is not a guarantee that issue content postdates model training.

How to read this score

Resolved-rate is useful when harness and budget are held constant; cross-system comparisons require matched conditions.

Similarity to real tasks: High — Tasks use real repository issues and tests, while agent harness, environment, and tool configuration still affect transfer.

Data contamination risk: High — The benchmark is based on public GitHub issues and repositories. Human verification improves task quality, but ModelRefs cannot rule out training overlap for any model run without source-scoped decontamination evidence.

Benchmark gaming risk: High — Scores depend on scaffold, tools, execution environment, cost or step budgets, and benchmark-specific agent tuning; cross-run comparisons require matched conditions.

What you still need to test yourself

  • Run representative issues from the target languages, repositories, dependencies, and security constraints.
  • Review patch maintainability, security, scope, and human acceptance beyond test passage.
  • Verify trace completeness, version attribution, tool and environment failures, regression sensitivity, incident linkage, human adjudication, and rollback behavior in the actual evaluation stack.

This benchmark supports decisions about:

  • Compare repository-level issue resolution under matched harness, tool, budget, and environment conditions.
  • Assess whether a coding system can coordinate changes across an existing codebase and pass benchmark tests.
  • Seed versioned coding-agent regression cohorts for an evaluation or observability workflow while keeping internal release criteria separate.

Limitations

  • Does not cover every language, repository type, security risk, or software-development workflow.
  • A resolved issue is not equivalent to production-ready code.
  • A benchmark delta does not prove that an evaluation pipeline or observability system detects, explains, or safely responds to production regressions.
  • Every score must identify the exact model, Verified subset or compatible subset, scaffold, tools, inference budget, and source; model training overlap remains unknown unless source-scoped evidence says otherwise.

Sources re-reviewed 2026-07-11. SWE-bench Verified is treated as an aging 2024 static subset; the 2024-08-13 release is provenance, not an issue-corpus freshness cutoff.

Sources

How this benchmark is scored

Categorycoding
Maximum score100 % resolved
DirectionHigher is better

Primary source: https://www.swebench.com/

Published results

ModelScoreRun dateSource
GPT-574.22026-05-01Aggregated public reports
Claude Opus 472.52026-05-01Aggregated public reports
DeepSeek R149.22026-05-01Aggregated public reports
GPT-5 Mini492026-05-01Aggregated public reports

Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to SWE-Bench Verified Leaderboard — AI Model Scores.