article
Multimodal Aspect-Based Sentiment Analysis aims to determine the sentiment polarity associated with specific aspects by jointly analyzing textual content and corresponding visual information. This fine-grained task plays a crucial role in applications such as brand monitoring, customer feedback analysis, and personalized recommendation systems. Despite recent progress in the field, existing models still suffer from two major limitations: suboptimal modality fusion and weak aspect-level alignment, both of which degrade overall performance. In this paper, we propose a novel attention-guided fusion framework designed two-stage attention mechanism. First, an attention-based fusion is performed between the aspect embeddings and the text embeddings to generate aspect-aware textual representations. These enriched textual features are then fused with visual embeddings using a secondary attention layer, enabling cross-modal interactions that are critical for sentiment interpretation. We evaluate our model on the MASAD multimodal benchmark dataset, where it consistently outperforms several state-of-the-art baselines across multiple categories. The results confirm the effectiveness of our approach in capturing aspect-specific sentiment cues by jointly leveraging textual and visual information in a unified and context-aware manner.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/icoa66896.2025.11236916
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.