TL;DR
Multimodal large language models (MLLMs) have not been thoroughly evaluated for their ability to mimic subjective human responses in video assessments. This study utilized MLLMs to act as synthetic participants in evaluating perceived sensory engagement with short videos, based on the Perceived Message Sensation Value framework.
✦ Why It Matters
Engineers and researchers can leverage MLLMs to streamline video engagement studies, enhancing efficiency and scalability.
Key Takeaways
How It Works
The study utilized the Perceived Message Sensation Value (PMSV) framework to measure subjective responses to videos, comparing human ratings with those generated by MLLMs. This involved analyzing a 17-item scale that assessed various emotional and sensory engagement metrics.
⚠ The Catch
MLLMs exhibited significant biases, including a downward mean-shift in ratings and inconsistent sensitivity to participant profiles, limiting their effectiveness as substitutes for human participants.
Related