ModelRefs / SWE-Bench Lite Leaderboard — AI Model Scores

SWE-Bench Lite Leaderboard — AI Model Scores

300-task subset of SWE-Bench for faster eval. Current leaders, methodology, and citation sources for SWE-Bench Lite.

Overview

300-task subset of SWE-Bench for faster eval.

How it is measured: Resolve-rate using agentless or agent pipelines.

How this benchmark is scored

Categorycoding
Maximum score100 % resolved
DirectionHigher is better

Primary source: https://www.swebench.com/lite.html

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to SWE-Bench Lite Leaderboard — AI Model Scores.