ModelRefs / RULER — AI Glossary
RULER — AI Glossary
A long-context benchmark with synthetic tasks at controlled context lengths to stress-test retrieval and reasoning across 4K–128K tokens.
Overview
RULER (Hsieh et al. 2024, NVIDIA) tests: needle in a haystack (single/multi), variable tracking, and aggregation tasks at lengths 4K, 8K, 16K, 32K, 64K, 128K. Unlike NIAH, RULER requires multi-hop retrieval and counting. Reveals that many models claiming 128K context show significant degradation beyond 32K on hard tasks.
Reference details
| Topic | evaluation |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Primary source
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to RULER — AI Glossary.
Frequently asked questions
What is RULER?
A long-context benchmark with synthetic tasks at controlled context lengths to stress-test retrieval and reasoning across 4K–128K tokens.
What concepts are related to RULER?
Closely related concepts include needle in haystack, long context, open llm leaderboard.