TL;DR
Evaluating the safety of large language models (LLMs) as judges is challenging due to their reliance on context and varying definitions of safety. This study investigates LLMs' susceptibility to context information and their adaptability to different safety standards.
✦ Why It Matters
Engineers should consider context and safety definition variability when using LLMs for safety evaluations.
Key Takeaways
Full Summary
Large language models (LLMs) are increasingly used as judges to assess safety in various applications, but their evaluation is often limited to simple benchmarks that measure agreement with human judgments. This research explores two critical aspects of LLMs-as-judges: their dependence on contextual information and their ability to adapt to different safety definitions, which may not match their inherent safety biases.
The study involved testing multiple generalist LLMs to assess their safety judgment capabilities under varying contexts. Findings reveal that LLMs frequently misalign with alternative safety definitions, suggesting that their internal safety priors can lead to inconsistent evaluations.
These insights emphasize the necessity for more nuanced evaluation frameworks that account for context and diverse safety perspectives. For engineers and researchers, this means rethinking how LLMs are assessed and potentially developing new methodologies for safety evaluation.
Related