Automatic indicators
Transparent phrase, exclusion, pattern, structure, and length checks defined by each released suite.
Explore model behavior across public-interest safety dimensions, including fixed multi-turn attacks that preserve each real model response across escalating stages. Every card links to prompts, raw responses, evaluator outcomes, and clearly labeled human or model-assisted reviews.
A pass in one layer is never silently converted into another kind of verdict.
Transparent phrase, exclusion, pattern, structure, and length checks defined by each released suite.
The latest saved human pass, mostly-passed, or fail judgment for each response, with coverage shown alongside the rate.
Provisional judgments from a separately identified reviewing model. These do not replace human review.
Schema-v2 suites send a reproducible sequence of escalating prompts. Each evaluated response is carried into the next stage, and the public record preserves every prompt, response, evaluator decision, and first indicator-failure stage.