ModelRefs / G-Eval — AI Glossary

G-Eval — AI Glossary

An LLM-as-judge framework that generates evaluation steps via chain-of-thought before scoring. The auto-generated rubric is a key feature.

Overview

G-Eval (Liu et al., 2023) outperforms vanilla LLM judging on summarization and dialogue tasks by forcing the judge to articulate explicit evaluation criteria before assigning a score. The auto-generated rubric is a key feature.

Reference details

Topicevaluation
Last reviewed2026-06-24

Primary source

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to G-Eval — AI Glossary.

Frequently asked questions

What is G-Eval?

An LLM-as-judge framework that generates evaluation steps via chain-of-thought before scoring.

What concepts are related to G-Eval?

Closely related concepts include llm as judge, eval, chain of thought.