TL;DR
OpenAI lacked comprehensive safety evaluations for GPT-5.1 variants, particularly around mental health risks and user emotional dependence. A system card addendum was created documenting updated safety metrics and new evaluation protocols for GPT-5.1 Instant and GPT-5.1 Thinking models.
✦ Why It Matters
Engineers can reference GPT-5.1 safety metrics to design guardrails and evaluate mental health risks in production deployments.
Key Takeaways
Full Summary
OpenAI released a system card addendum—a safety documentation supplement—for GPT-5.1 Instant and GPT-5.1 Thinking, two variants of their large language model. The addendum addresses a gap in existing safety evaluations by introducing new assessment categories focused on mental health impacts and emotional reliance (user over-dependence on AI for emotional support).
The evaluation methodology includes structured safety metrics designed to measure potential harms in these specific domains. Results quantify baseline safety performance across both model variants, enabling comparison and risk stratification.
This work establishes precedent for domain-specific safety evaluation in frontier AI systems and provides engineers with concrete benchmarks for assessing model behavior in sensitive use cases.
Related