ModelRefs / Safety Benchmarks — Top AI Models

Safety Benchmarks — Top AI Models

Red-team, jailbreak, toxicity, and refusal-quality evaluations. Quantify resistance to misuse and harmful-output rate.

Overview

Red-team, jailbreak, toxicity, and refusal-quality evaluations.

What this category is for: Quantify resistance to misuse and harmful-output rate.

Benchmarks in this category

  • HarmBench — Standardized red-teaming eval covering 510 harmful behaviors.
  • JailbreakBench — Reproducible jailbreak evaluation across 100 misuse behaviors.
  • ToxiGen — Implicit toxicity classification across 13 demographic groups.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Safety Benchmarks — Top AI Models.