ModelRefs / LongBench v2 Leaderboard — AI Model Scores

LongBench v2 Leaderboard — AI Model Scores

Realistic long-context tasks up to 2M tokens. Current leaders, methodology, and citation sources for LongBench v2.

Overview

Realistic long-context tasks up to 2M tokens.

How it is measured: Macro accuracy across 503 multiple-choice items.

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://longbench2.github.io/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to LongBench v2 Leaderboard — AI Model Scores.