ModelRefs / Best Open-Source Models
Best Open-Source Models
The most capable open-weight models you can self-host today.
Overview
The most capable open-weight models you can self-host today.
How this ranking is produced
24 models in the ModelRefs catalogue carry qualifying benchmark evidence for this category. The ten highest-scoring are listed below.
Scores below are a heuristic over the benchmark evidence ModelRefs holds for each model, not a guarantee of real-world performance. A model ranks only where it has qualifying benchmark results, so a capable model with thin evidence can rank low or be absent. Each entry states the benchmarks behind its score — read those before acting on the order.
Ranked models
-
#1 Llama 3.1 405B
Score 139 out of 100. Open-weight model (Llama 3.1 Community) with benchmark avg 88.6.
-
#2 Phi-4
Score 134 out of 100. Open-weight model (MIT) with benchmark avg 83.7.
-
#3 Llama 3.1 70B
Score 134 out of 100. Open-weight model (Llama 3.1 Community) with benchmark avg 83.6.
-
#4 Nemotron-4 340B
Score 129 out of 100. Open-weight model (NVIDIA Open Model) with benchmark avg 78.7.
-
#5 DeepSeek V3
Score 126 out of 100. Open-weight model (MIT) with benchmark avg 75.9.
-
#6 Command R+
Score 126 out of 100. Open-weight model (CC-BY-NC) with benchmark avg 75.7.
-
#7 Llama 4 Scout
Score 124 out of 100. Open-weight model (Llama 4 Community) with benchmark avg 74.3.
-
#8 Qwen 2.5 72B
Score 121 out of 100. Open-weight model (Qwen License) with benchmark avg 71.1.
-
#9 Llama 3.1 8B
Score 119 out of 100. Open-weight model (Llama 3.1 Community) with benchmark avg 69.4.
-
#10 Mistral Nemo
Score 118 out of 100. Open-weight model (Apache-2.0) with benchmark avg 68.0.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Open-Source Models.