ModelRefs / Refusal — AI Glossary

Refusal — AI Glossary

A model declining to respond to a request, citing safety, policy, capability, or ethical reasons. Also called rejection or model refusal.

Overview

Refusals are explicit safety behaviors trained into models to prevent harmful content generation. 'Over-refusal' (refusing benign requests) and 'under-refusal' (complying with harmful requests) are dual failure modes measured by safety benchmarks. Refusal calibration is a key alignment challenge—Claude's Constitutional AI, OpenAI's usage policies, and Anthropic's RSP all govern refusal thresholds.

Reference details

Topicsafety
Also known asrejection, model refusal, content refusal
Last reviewed2026-06-24

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Refusal — AI Glossary.

Frequently asked questions

What is Refusal?

A model declining to respond to a request, citing safety, policy, capability, or ethical reasons.

Is Refusal the same as rejection?

Yes — rejection, model refusal, content refusal are common aliases for Refusal.

What concepts are related to Refusal?

Closely related concepts include guardrails, alignment, abstention.