TL;DR
Omni-Multimodal Large Language Models (Omni-MLLMs) struggle with audio-visual intelligence due to a lack of evaluation benchmarks. AVI-Bench was developed to assess these models through three stages: perception, understanding, and reasoning, using cross-modal tasks.
✦ Why It Matters
Engineers can use AVI-Bench to evaluate and improve the audio-visual capabilities of their Omni-MLLMs.
Key Takeaways
How It Works
AVI-Bench evaluates models through a series of tasks that require them to interpret and integrate audio and visual information simultaneously. The benchmark is structured into three stages: perception, understanding, and reasoning, allowing for a comprehensive assessment of model capabilities.
The extension, AVI-Bench-PriSe, introduces low-semantic stimuli to test models' primitive audio-visual sensations, pushing their limits beyond familiar training data.
Related