ModelRefs / Best Coding Models
Best Coding Models
Top AI models ranked by code-generation benchmarks like HumanEval and MBPP.
Overview
Top AI models ranked by code-generation benchmarks like HumanEval and MBPP.
How this ranking is produced
18 models in the ModelRefs catalogue carry qualifying benchmark evidence for this category. The ten highest-scoring are listed below.
Scores below are a heuristic over the benchmark evidence ModelRefs holds for each model, not a guarantee of real-world performance. A model ranks only where it has qualifying benchmark results, so a capable model with thin evidence can rank low or be absent. Each entry states the benchmarks behind its score — read those before acting on the order.
Ranked models
-
#1 Phi-4
Score 83 out of 100. HumanEval/MBPP avg 82.6 with MMLU support of 84.8.
-
#2 DeepSeek R1
Score 46 out of 100. HumanEval/MBPP avg 65.9 with MMLU support of 0.0.
-
#3 Llama 4 Scout
Score 45 out of 100. HumanEval/MBPP avg 32.8 with MMLU support of 74.3.
-
#4 QwQ 32B
Score 44 out of 100. HumanEval/MBPP avg 63.4 with MMLU support of 0.0.
-
#5 Llama 3.1 405B
Score 27 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 88.6.
-
#6 GPT-4o
Score 26 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 87.2.
-
#7 Gemini 1.5 Pro
Score 26 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 85.9.
-
#8 Llama 3.1 70B
Score 25 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 83.6.
-
#9 Nemotron-4 340B
Score 24 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 78.7.
-
#10 Gemini 2.0 Flash
Score 23 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 77.6.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Coding Models.