ModelRefs / Constitutional AI (CAI) — AI Glossary

Constitutional AI (CAI) — AI Glossary

Anthropic's alignment method where a model critiques and revises its own outputs against a written set of principles. Powers Claude's safety training.

Overview

CAI reduces dependence on human red-teamers for harmlessness training. The model is first trained to write critiques against the constitution, then to revise outputs. Successor to supervised harmlessness filtering. Powers Claude's safety training.

Reference details

Topicsafety
Last reviewed2026-06-24

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Constitutional AI (CAI) — AI Glossary.

Frequently asked questions

What is Constitutional AI (CAI)?

Anthropic's alignment method where a model critiques and revises its own outputs against a written set of principles.

What concepts are related to Constitutional AI (CAI)?

Closely related concepts include alignment, rlhf, red teaming.