TL;DR
Large language models (LLMs) show stable self-reports on personality traits, but these do not predict actual behavior. A new psychometric instrument was developed based on LLM behaviors, using exploratory factor analysis to identify five key dimensions.
✦ Why It Matters
Engineers and researchers should be cautious when using LLM self-reports for behavioral predictions in applications.
Key Takeaways
Full Summary
Self-reported personality traits from large language models (LLMs) have been found to be stable but not predictive of actual behavior, raising questions about the alignment of LLMs with human personality constructs. To address this, a novel psychometric instrument was created, derived from LLM behavioral affordances through exploratory factor analysis (EFA).
The instrument included 300 items across 12 behavioral dimensions and identified five factors: Responsiveness, Deference, Boldness, Guardedness, and Verbosity, demonstrating excellent reliability. However, when comparing self-reports to 2,500 behavioral samples rated by humans and LLM judges, self-reports showed negligible correlation with human ratings.
Notably, while LLM judges agreed with human ratings, self-reports did not align, suggesting a unique variance shared between LLM self-reports and LLM judges that humans do not perceive. This research highlights the limitations of LLM self-assessment and its implications for using LLMs in evaluative roles.
Related