ModelRefs / BigCodeBench Leaderboard — AI Model Scores

BigCodeBench Leaderboard — AI Model Scores

Function-level synthesis requiring practical library use. Current leaders, methodology, and citation sources for BigCodeBench.

Overview

Function-level synthesis requiring practical library use.

How it is measured: Pass@1 across complete + instruct splits.

How this benchmark is scored

Categorycoding
Maximum score100 pass@1
DirectionHigher is better

Primary source: https://bigcode-bench.github.io/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to BigCodeBench Leaderboard — AI Model Scores.