MARATTO

article

Sexism Discovery using CNN, Word Embeddings, NLP and Data Augmentation

Abstract

The pervasive issue of online sexism continues to pose significant challenges, fostering environments characterized by toxicity and perpetuating harmful societal norms. In response, this paper presents an approach for the discovery of sexist statements employing convolutional neural networks (CNNs), Word Embeddings, and data augmentation techniques. Through the fusion of CNNs’ capacity for hierarchical feature extraction with the semantic representations afforded by Word Embeddings, our method achieves exemplary discrimination performance. Additionally, the incorporation of data augmentation enriches the training dataset, thereby augmenting model generalization and resilience. Empirical evaluation on a larger dataset of statements demonstrates the efficacy of our approach, surpassing many baseline approaches in terms of discovery accuracy, precision, recall and F1-score.

Research topics

  • Hate Speech and Cyberbullying Detection
  • Computational and Text Analysis Methods

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/codit62066.2024.10708284

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.