ModelRefs / BrowseComp Long Context Leaderboard — AI Model Scores

BrowseComp Long Context Leaderboard — AI Model Scores

OpenAI long-context question-answering benchmark with relevant search results embedded in inputs up to hundreds of thousands of tokens. Current leaders, methodology, and citation sources for BrowseComp Long Context.

Overview

OpenAI long-context question-answering benchmark with relevant search results embedded in inputs up to hundreds of thousands of tokens.

How it is measured: Exact-answer accuracy over grounded long-context question-answering tasks, reported separately by input-length range.

How this benchmark is scored

Categoryretrieval
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://huggingface.co/datasets/openai/browsecomp-long-context

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to BrowseComp Long Context Leaderboard — AI Model Scores.