ModelRefs / MLVU Leaderboard — AI Model Scores

MLVU Leaderboard — AI Model Scores

Multi-task long video understanding (3 min – 2 hr clips). Current leaders, methodology, and citation sources for MLVU.

Overview

Multi-task long video understanding (3 min – 2 hr clips).

How it is measured: M-Avg score across 9 tasks.

How this benchmark is scored

Categorymultimodal
Maximum score100 M-Avg
DirectionHigher is better

Primary source: https://github.com/JUNJIE99/MLVU

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MLVU Leaderboard — AI Model Scores.