ModelRefs / MultiPL-E Leaderboard — AI Model Scores

MultiPL-E Leaderboard — AI Model Scores

HumanEval & MBPP translated into 18 programming languages. Current leaders, methodology, and citation sources for MultiPL-E.

Overview

HumanEval & MBPP translated into 18 programming languages.

How it is measured: pass@1 averaged across all languages.

How this benchmark is scored

Categorycoding
Maximum score100 pass@1
DirectionHigher is better

Primary source: https://nuprl.github.io/MultiPL-E/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MultiPL-E Leaderboard — AI Model Scores.