ModelRefs / MBPP+ Leaderboard — AI Model Scores
MBPP+ Leaderboard — AI Model Scores
Mostly Basic Python Problems with adversarial extended tests. Current leaders, methodology, and citation sources for MBPP+.
Overview
Mostly Basic Python Problems with adversarial extended tests.
How it is measured: pass@1 with EvalPlus hidden tests.
How this benchmark is scored
| Category | coding |
|---|---|
| Maximum score | 100 pass@1 |
| Direction | Higher is better |
Primary source: https://github.com/evalplus/evalplus
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MBPP+ Leaderboard — AI Model Scores.