ModelRefs / SWE-Bench Verified — AI Glossary

SWE-Bench Verified — AI Glossary

A 500-instance human-validated subset of SWE-Bench with confirmed solvable GitHub issues, the standard subset for agent benchmarking.

Overview

SWE-Bench Verified (OpenAI, Anthropic 2024) curates 500 SWE-Bench instances confirmed by human annotators to be solvable and unambiguously specified. Removes noise that caused low correlation between human judgment and automated grading in the full 2,294-instance set. Frontier agents (Claude 3.5 Sonnet, SWE-agent) score 40–55% on Verified.

Reference details

Topicevaluation
Last reviewed2026-06-24

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to SWE-Bench Verified — AI Glossary.

Frequently asked questions

What is SWE-Bench Verified?

A 500-instance human-validated subset of SWE-Bench with confirmed solvable GitHub issues, the standard subset for agent benchmarking.

What concepts are related to SWE-Bench Verified?

Closely related concepts include coding agent, livecodebench, aider polyglot.