ModelRefs / Best Coding Models

Best Coding Models

Top AI models ranked by code-generation benchmarks like HumanEval and MBPP.

Overview

Top AI models ranked by code-generation benchmarks like HumanEval and MBPP.

How this ranking is produced

18 models in the ModelRefs catalogue carry qualifying benchmark evidence for this category. The ten highest-scoring are listed below.

Scores below are a heuristic over the benchmark evidence ModelRefs holds for each model, not a guarantee of real-world performance. A model ranks only where it has qualifying benchmark results, so a capable model with thin evidence can rank low or be absent. Each entry states the benchmarks behind its score — read those before acting on the order.

Ranked models

  1. #1 Phi-4

    Score 83 out of 100. HumanEval/MBPP avg 82.6 with MMLU support of 84.8.

  2. #2 DeepSeek R1

    Score 46 out of 100. HumanEval/MBPP avg 65.9 with MMLU support of 0.0.

  3. #3 Llama 4 Scout

    Score 45 out of 100. HumanEval/MBPP avg 32.8 with MMLU support of 74.3.

  4. #4 QwQ 32B

    Score 44 out of 100. HumanEval/MBPP avg 63.4 with MMLU support of 0.0.

  5. #5 Llama 3.1 405B

    Score 27 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 88.6.

  6. #6 GPT-4o

    Score 26 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 87.2.

  7. #7 Gemini 1.5 Pro

    Score 26 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 85.9.

  8. #8 Llama 3.1 70B

    Score 25 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 83.6.

  9. #9 Nemotron-4 340B

    Score 24 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 78.7.

  10. #10 Gemini 2.0 Flash

    Score 23 out of 100. HumanEval/MBPP avg 0.0 with MMLU support of 77.6.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Coding Models.