TL;DR
Large language models improve with preference alignment, but internal changes remain unclear. MENTIS, a geometry-first framework, was developed to measure these changes using torsion norms and energy metrics.
✦ Why It Matters
Engineers can leverage MENTIS to better understand and improve the internal workings of aligned language models.
Key Takeaways
Full Summary
Preference alignment enhances the behavior of large language models, yet it is uncertain what internal changes occur during this process. MENTIS is introduced as a framework to analyze the geometric structure of language models before and after alignment, focusing on layerwise covariance-based torsion norms (T1), spectral torsion diagnostics (T2), and an Energy-Radiance-Activation measure (ERA).
The research involved four model pairs, revealing that alignment-induced changes are not uniform; normative concepts showed larger torsion shifts compared to factual ones. Additionally, torsion was negatively correlated with contextual entropy, indicating that more structured changes occur in specific layers, particularly mid-to-late layers of the architecture.
These findings suggest that preference alignment leaves measurable geometric signatures in internal computations, which behavior-level evaluations alone cannot capture.
Related