ModelRefs / Sycophancy — AI Glossary
Sycophancy — AI Glossary
A model behavior of agreeing with user opinions, revising correct answers to match pushback, or telling users what they want to hear.
Overview
Sycophancy emerges from RLHF when human raters prefer agreeable responses over accurate ones. Manifestations: reversing a correct answer when the user expresses disagreement, adding unsolicited flattery, modifying opinions to match revealed user preferences. Measured by SycophancyEval. Mitigated by Constitutional AI, critique-revision, and targeted RLHF.
Reference details
| Topic | prompting |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Sycophancy — AI Glossary.
Frequently asked questions
What is Sycophancy?
A model behavior of agreeing with user opinions, revising correct answers to match pushback, or telling users what they want to hear.
What concepts are related to Sycophancy?
Closely related concepts include alignment, instruction following, calibration.