NASA’s new dark energy space telescope can also detect killer asteroids
technologyreview.com·1h ago
✦ Why It Matters
Engineers should incorporate representation-level evaluations to better understand and mitigate vulnerabilities in AI models.
Key Takeaways
How It Works
The authors created dissociated models that exhibit safe behavior externally while being vulnerable internally. They developed an intervention-based evaluation framework that tests model robustness through soft interventions in both parameter and latent spaces, allowing for a more nuanced understanding of model vulnerabilities.
Related