MARATTO

article · Revue d intelligence artificielle

Towards Amazigh Word Embedding: Corpus Creation and Word2Vec Models Evaluations

20233 citationsOpen accessUniversité Sultan Moulay Slimane

Abstract

Distributed representations of words in a vector space help learning algorithms to model semantic notions of word similarity and distances in sentences.Most of the existing researches have been done on the Latin, Arabic, and other language, while the Amazigh language is ignored.In this paper, we try to build a first model word embeddings for Amazigh language and describe the steps needed to build it.Therefore, we implement a Word2Vec that a combination of two techniques -CBOW (Continuous bag of words) and Skip-gram to transform words written in Tifinagh to vector form.To obtain the highest performance, we evaluate two parameters of Word2Vec include Word2Vec model architecture and vector dimension.This evaluation process was implemented towards our proposed corpus collected on Amazigh websites for different domains.The result shows that the highest accuracy values are obtained under the combination of CBOW model and 300 dimensional vector.

Research topics

  • Natural Language Processing Techniques
  • Text Readability and Simplification
  • Topic Modeling

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.18280/ria.370324

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.