ModelRefs / MT-Bench — AI Glossary

MT-Bench — AI Glossary

A multi-turn chat evaluation benchmark using GPT-4 as judge to rate models on 80 two-turn questions across 8 domains. Scores range 1–10.

Overview

MT-Bench (Zheng et al. 2023, LMSYS) tests instruction following in multi-turn conversations across writing, roleplay, extraction, reasoning, math, coding, STEM, and humanities. GPT-4 scores as a judge correlate well with human preferences. Used alongside Chatbot Arena to rank chat models. Scores range 1–10.

Reference details

Topicevaluation
Last reviewed2026-06-24

Primary source

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MT-Bench — AI Glossary.

Frequently asked questions

What is MT-Bench?

A multi-turn chat evaluation benchmark using GPT-4 as judge to rate models on 80 two-turn questions across 8 domains.

What concepts are related to MT-Bench?

Closely related concepts include alpacaeval, chatbot arena, lmsys.