ModelRefs / Best Multimodal Models

Best Multimodal Models

Top vision + language models for image, audio and text understanding.

Overview

Top vision + language models for image, audio and text understanding.

How this ranking is produced

25 models in the ModelRefs catalogue carry qualifying benchmark evidence for this category. The ten highest-scoring are listed below.

Scores below are a heuristic over the benchmark evidence ModelRefs holds for each model, not a guarantee of real-world performance. A model ranks only where it has qualifying benchmark results, so a capable model with thin evidence can rank low or be absent. Each entry states the benchmarks behind its score — read those before acting on the order.

Ranked models

  1. #1 Qwen2.5-VL 72B

    Score 133 out of 100. Native multimodal model with 83.3 multimodal benchmark.

  2. #2 Claude Opus 4

    Score 127 out of 100. Native multimodal model with 76.5 multimodal benchmark.

  3. #3 Llama 4 Scout

    Score 126 out of 100. Native multimodal model with 75.8 multimodal benchmark.

  4. #4 Claude Sonnet 4

    Score 124 out of 100. Native multimodal model with 74.4 multimodal benchmark.

  5. #5 Pixtral 12B

    Score 122 out of 100. Native multimodal model with 71.6 multimodal benchmark.

  6. #6 InternVL2.5 78B

    Score 120 out of 100. Native multimodal model with 70.1 multimodal benchmark.

  7. #7 GPT-5

    Score 50 out of 100. Native multimodal model.

  8. #8 GPT-5 Mini

    Score 50 out of 100. Native multimodal model.

  9. #9 o3

    Score 50 out of 100. Native multimodal model.

  10. #10 o4 Mini

    Score 50 out of 100. Native multimodal model.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Multimodal Models.