AI Safety Research Areas

Public-interest AI safety research areas in development. Formal reports, code, and supporting materials will be linked here as they are released.

Current Work in Development

These areas describe active and planned work. They are not presented as peer-reviewed publications.

AlignmentIn development

Value Consistency Evaluation

Developing a framework for studying whether model outputs remain consistent with stated human values across varied prompting conditions.

EvaluationIn development

Red-Teaming Benchmark Methods

Reviewing model red-teaming approaches and designing repeatable evaluation methods for public-interest safety work.

PolicyIn development

AI Safety Thresholds

Studying how public safety thresholds could help organizations reason about responsible deployment and risk escalation.

RobustnessIn development

Specification Gaming Patterns

Cataloging specification-gaming patterns and mitigation ideas for future educational and research materials.

PolicyIn development

AI Incident Reporting

Developing public guidance for documenting, triaging, and learning from AI safety incidents.

AlignmentIn development

Corrigibility and Oversight

Studying how human oversight and correction mechanisms can degrade under new contexts and long-horizon tasks.

SecurityIn development

Prompt Injection Risk

Developing educational material and evaluation approaches for prompt injection risks in LLM-integrated systems.

SecurityIn development

Information Poisoning

Studying threats to training data, retrieval systems, fine-tuning workflows, and agent memory.