ModelRefs / MBPP+ Leaderboard — AI Model Scores

MBPP+ Leaderboard — AI Model Scores

Mostly Basic Python Problems with adversarial extended tests. Current leaders, methodology, and citation sources for MBPP+.

Overview

Mostly Basic Python Problems with adversarial extended tests.

How it is measured: pass@1 with EvalPlus hidden tests.

How this benchmark is scored

Categorycoding
Maximum score100 pass@1
DirectionHigher is better

Primary source: https://github.com/evalplus/evalplus

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MBPP+ Leaderboard — AI Model Scores.