TL;DR
Multimodal large language models (LLMs) struggle to develop efficient communication strategies in repeated reference games, failing to create partner-specific conventions. This study investigates whether their alignment on labels is due to shared vocabulary rather than specific grounding with partners.
✦ Why It Matters
Engineers can focus on improving multimodal LLMs to enhance their adaptability in communication tasks.
Key Takeaways
How It Works
The study employs a constrained pseudo-dyad baseline to isolate the effects of partner history on label alignment. By comparing the communication strategies of multimodal LLMs and humans, it reveals that LLMs rely on fixed verbosity rather than adapting their language based on shared experiences.
Related