ModelRefs / MT-Bench Leaderboard — AI Model Scores

MT-Bench Leaderboard — AI Model Scores

LMSYS multi-turn chatbot benchmark across 8 categories, GPT-4 graded 1–10. Current leaders, methodology, and citation sources for MT-Bench.

Overview

LMSYS multi-turn chatbot benchmark across 8 categories, GPT-4 graded 1–10.

How it is measured: 80 two-turn questions; mean GPT-4 score across writing, math, coding, reasoning, extraction, STEM, humanities, roleplay.

What this benchmark measures

  • multi-turn instruction following
  • open-ended chat response quality
  • conversation consistency

Relevant to:

  • conversational-assistant screening
  • reviewable customer-communication drafting evaluation
  • limited conversation-quality screening for matter intake, payer-policy questions, and investor-draft review

Failure modes it exercises:

  • weak follow-up response
  • instruction loss across turns
  • low judged helpfulness or relevance

Method and its limits

LLM-as-a-judge scoring of responses to a fixed set of multi-turn open-ended questions, interpreted with the exact question set, generation settings, judge model, judge prompt, reference-answer use, and aggregation method.

  • The primary paper documents position, verbosity, self-enhancement, and reasoning biases in LLM judges.
  • Judge model, prompt, reference answers, response order, sampling, and implementation revision can materially affect scores.

Dataset

Dataset
MT-Bench
Type
fixed multi-turn open-ended question set evaluated by an LLM judge
Freshness
unknown

The primary sources establish a broad multi-turn question set, not current-world, collections, finance, compliance, or organization-specific communication coverage.

How to read this score

Treat the score as judge- and protocol-scoped evidence of broad chat quality; inspect per-category outputs and human review rather than using it as a compliance or communication-approval signal.

Similarity to real tasks: Medium — Multi-turn open-ended prompts resemble conversational interaction, but the questions and judge rubric do not reproduce account grounding, consent, policy, finance, or safety constraints.

Data contamination risk: Unknown — Questions and evaluation code are public, and no model-run-specific overlap assessment is registered.

Benchmark gaming risk: Unknown — Public questions, judge preferences, style effects, and benchmark-targeted tuning create unresolved optimization risk.

What you still need to test yourself

  • Test grounded customer communications using representative account facts, approved templates, tone rules, consent states, disputes, and escalation cases.
  • Use qualified human reviewers to assess factuality, misleading language, privacy, policy adherence, safety, and jurisdiction-specific requirements.
  • For legal, clinical, payer, or investor workflows, test source-grounded conversations with confidentiality states, professional handoffs, abstention, citation validity, and prohibited-advice or approval boundaries.

This benchmark supports decisions about:

  • Screen conversational systems under the same MT-Bench question, generation, judge, and aggregation protocol.
  • Identify multi-turn and response-quality weaknesses that warrant deeper human-reviewed communication tests.

Limitations

  • MT-Bench is not evidence that a message is factually grounded, legally permitted, policy-compliant, safe to send, or effective for collections.
  • Model-judged preference can reward style and verbosity and must be calibrated against human review on the target workflow.
  • It does not establish legal-intake safety, clinical or coverage validity, investor disclosure compliance, or readiness for a consequential conversation.

Sources reviewed 2026-07-02. Dataset freshness remains unknown; pin the question set, repository revision, generation settings, judge model, judge prompt, references, ordering, and aggregation.

Sources

How this benchmark is scored

Categoryopen-source
Maximum score10 score /10
DirectionHigher is better

Primary source: https://huggingface.co/spaces/lmsys/mt-bench

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MT-Bench Leaderboard — AI Model Scores.