TL;DR
Multimodal large language models (MLLMs) have not been thoroughly evaluated for their ability to mimic subjective human responses in video assessments. This study utilized MLLMs to act as synthetic participants in evaluating perceived sensory engagement with short videos, based on the Perceived Message Sensation Value framework.
✦ Why It Matters
Engineers and researchers can leverage MLLMs to streamline video engagement studies, enhancing efficiency and scalability.
Key Takeaways
Full Summary
Multimodal large language models (MLLMs) integrate various data types, such as text and video, to perform tasks like understanding and reasoning. This research aimed to evaluate MLLMs as synthetic participants in assessing perceived sensory engagement, which refers to how engaging a video feels to viewers based on their subjective experiences.
Using the Perceived Message Sensation Value framework, the study compared MLLM outputs with human responses to short videos. The methodology involved training MLLMs on video content and measuring their responses against human evaluations.
Findings revealed that MLLMs could closely approximate human perceptions, achieving a correlation coefficient of 0.85 with human ratings. These results imply that MLLMs can be effectively used in video-based studies, potentially reducing the need for extensive human participant involvement.
Related