article · Procedia Computer Science
With the growth in the use of email and social media, spam has become a major challenge. With the rise of multimedia technologies, the prevalence of multimodal spam containing a mixture of text and images has significantly increased. However, most of the methods proposed to detect spam in the past are mainly based on text analysis. The development of a multimodal approach to spam filtering is therefore of paramount importance. The paper aims to develop an improved method for multimedia spam detection using hierarchical and multimodal message analysis, combined with deep learning (DL). Our approach is based on extracting several representative characteristics from multimodal data (text, links, images) using models based on large language models and convolutional networks. This aims to obtain a fine-grained semantic representation with focus on the key elements of the messages for more effective classification of multimedia spam. We experimented our method on a large corpus and used qualitative and quantitative analysis to compare accuracy and robustness. The experiments demonstrate the model’s capability to detect effectively complex spam strategies, reflecting significant improvement in spam identifying techniques.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1016/j.procs.2025.09.513
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.