ModelRefs / Fastest Models

Fastest Models

Lowest-latency models for real-time and high-throughput workloads.

Overview

Lowest-latency models for real-time and high-throughput workloads.

How this ranking is produced

44 models in the ModelRefs catalogue carry qualifying benchmark evidence for this category. The ten highest-scoring are listed below.

Scores below are a heuristic over the benchmark evidence ModelRefs holds for each model, not a guarantee of real-world performance. A model ranks only where it has qualifying benchmark results, so a capable model with thin evidence can rank low or be absent. Each entry states the benchmarks behind its score — read those before acting on the order.

Ranked models

  1. #1 GPT-5

    Score 50 out of 100. Estimated speed based on —.

  2. #2 GPT-5 Mini

    Score 50 out of 100. Estimated speed based on —.

  3. #3 Claude Opus 4

    Score 50 out of 100. Estimated speed based on —.

  4. #4 Llama 4 Scout

    Score 50 out of 100. Estimated speed based on —.

  5. #5 DeepSeek R1

    Score 50 out of 100. Estimated speed based on —.

  6. #6 Mistral Large 2

    Score 50 out of 100. Estimated speed based on —.

  7. #7 Command R+

    Score 50 out of 100. Estimated speed based on —.

  8. #8 o3

    Score 50 out of 100. Estimated speed based on —.

  9. #9 o4 Mini

    Score 50 out of 100. Estimated speed based on —.

  10. #10 Text Embedding 3 Large

    Score 50 out of 100. Estimated speed based on —.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Fastest Models.