TL;DR
Multimodal attention-based models have advanced significantly, but their decision-making processes remain largely opaque. A systematic review of literature from 2020 to 2024 was conducted to assess explainability techniques in these models.
✦ Why It Matters
Engineers can enhance their multimodal AI systems by adopting standardized evaluation practices for explainability.
Key Takeaways
Full Summary
Recent advancements in multimodal learning, which combines different types of data (like text and images), have been driven by attention-based models that enhance performance across various tasks. This systematic review analyzed research published between January 2020 and early 2024, focusing on explainability in these models.
It highlighted that most studies concentrate on vision-language and language-only models, primarily using attention mechanisms for explanations. However, these methods often fail to fully capture the interactions between different data types, and evaluation methods for explainability are inconsistent and lack robustness.
The review synthesizes findings and suggests a comprehensive set of recommendations to improve evaluation and reporting practices in multimodal XAI research. By addressing these gaps, the goal is to foster more interpretable and accountable AI systems that prioritize explainability.
Related