ModelRefs / CRUXEval Leaderboard — AI Model Scores

CRUXEval Leaderboard — AI Model Scores

Code reasoning over input prediction and output prediction. Current leaders, methodology, and citation sources for CRUXEval.

Overview

Code reasoning over input prediction and output prediction.

How it is measured: Pass@1 averaged across input and output tasks.

How this benchmark is scored

Categorycoding
Maximum score100 pass@1
DirectionHigher is better

Primary source: https://crux-eval.github.io/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to CRUXEval Leaderboard — AI Model Scores.