ModelRefs / Red Teaming — AI Glossary

Red Teaming — AI Glossary

Adversarial testing of an AI system by humans (or other models) attempting to elicit harmful, biased, or off-policy outputs.

Overview

Red teaming is required practice for frontier model releases and a core EU AI Act / NIST AI RMF expectation. Automated red teaming uses adversarial LLMs to scale coverage. Findings feed safety training and guardrail updates.

Reference details

Topicsafety
Last reviewed2026-06-24

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Red Teaming — AI Glossary.

Frequently asked questions

What is Red Teaming?

Adversarial testing of an AI system by humans (or other models) attempting to elicit harmful, biased, or off-policy outputs.

What concepts are related to Red Teaming?

Closely related concepts include jailbreak, guardrails, alignment.