MARATTO

article · Procedia Computer Science

A morphological word embedding for Arabic language

20241 citationOpen accessMohamed I University

Abstract

Word embedding is a technique for representing words as dense real vectors and is crucial for most natural language processing tasks. Common approaches include non-contextual embeddings such as Word2vec and contextual embeddings like BERT, both of which operate on the principle that words appearing in similar contexts have close vector representations. Despite their effectiveness, these methods often fail to capture some morphological information useful in many natural language processing applications. In this study, we propose an original word embedding of Arabic words based on the morphological characteristics of the words. Tests show that integrating this morphology-based word embedding into a topic detection model outperforms models using Word2vec or AraBERT embeddings.

Research topics

  • Topic Modeling
  • Natural Language Processing Techniques
  • Advanced Text Analysis Techniques

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1016/j.procs.2024.10.190

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.