ModelRefs / MT-Bench Leaderboard — AI Model Scores
MT-Bench Leaderboard — AI Model Scores
LMSYS multi-turn chatbot benchmark across 8 categories, GPT-4 graded 1–10. Current leaders, methodology, and citation sources for MT-Bench.
Overview
LMSYS multi-turn chatbot benchmark across 8 categories, GPT-4 graded 1–10.
How it is measured: 80 two-turn questions; mean GPT-4 score across writing, math, coding, reasoning, extraction, STEM, humanities, roleplay.
What this benchmark measures
- multi-turn instruction following
- open-ended chat response quality
- conversation consistency
Relevant to:
- conversational-assistant screening
- reviewable customer-communication drafting evaluation
- limited conversation-quality screening for matter intake, payer-policy questions, and investor-draft review
Failure modes it exercises:
- weak follow-up response
- instruction loss across turns
- low judged helpfulness or relevance
Method and its limits
LLM-as-a-judge scoring of responses to a fixed set of multi-turn open-ended questions, interpreted with the exact question set, generation settings, judge model, judge prompt, reference-answer use, and aggregation method.
- The primary paper documents position, verbosity, self-enhancement, and reasoning biases in LLM judges.
- Judge model, prompt, reference answers, response order, sampling, and implementation revision can materially affect scores.
Dataset
- Dataset
- MT-Bench
- Type
- fixed multi-turn open-ended question set evaluated by an LLM judge
- Freshness
- unknown
The primary sources establish a broad multi-turn question set, not current-world, collections, finance, compliance, or organization-specific communication coverage.
How to read this score
Treat the score as judge- and protocol-scoped evidence of broad chat quality; inspect per-category outputs and human review rather than using it as a compliance or communication-approval signal.
Similarity to real tasks: Medium — Multi-turn open-ended prompts resemble conversational interaction, but the questions and judge rubric do not reproduce account grounding, consent, policy, finance, or safety constraints.
Data contamination risk: Unknown — Questions and evaluation code are public, and no model-run-specific overlap assessment is registered.
Benchmark gaming risk: Unknown — Public questions, judge preferences, style effects, and benchmark-targeted tuning create unresolved optimization risk.
What you still need to test yourself
- Test grounded customer communications using representative account facts, approved templates, tone rules, consent states, disputes, and escalation cases.
- Use qualified human reviewers to assess factuality, misleading language, privacy, policy adherence, safety, and jurisdiction-specific requirements.
- For legal, clinical, payer, or investor workflows, test source-grounded conversations with confidentiality states, professional handoffs, abstention, citation validity, and prohibited-advice or approval boundaries.
This benchmark supports decisions about:
- Screen conversational systems under the same MT-Bench question, generation, judge, and aggregation protocol.
- Identify multi-turn and response-quality weaknesses that warrant deeper human-reviewed communication tests.
Limitations
- MT-Bench is not evidence that a message is factually grounded, legally permitted, policy-compliant, safe to send, or effective for collections.
- Model-judged preference can reward style and verbosity and must be calibrated against human review on the target workflow.
- It does not establish legal-intake safety, clinical or coverage validity, investor disclosure compliance, or readiness for a consequential conversation.
Sources reviewed 2026-07-02. Dataset freshness remains unknown; pin the question set, repository revision, generation settings, judge model, judge prompt, references, ordering, and aggregation.
Sources
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena LMSYS / research collaborators · accessed 2026-07-02
- FastChat MT-Bench and LLM-judge instructions LMSYS · accessed 2026-07-02
How this benchmark is scored
| Category | open-source |
|---|---|
| Maximum score | 10 score /10 |
| Direction | Higher is better |
Primary source: https://huggingface.co/spaces/lmsys/mt-bench
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MT-Bench Leaderboard — AI Model Scores.