MARATTO

conference paper

Adapting a decision tree based tagger for Arabic

Abstract

Several probabilistic methods used for Part of speech (POS) tagging are based on Hidden Markov Models (HMM), these methods have difficulties especially in estimating transition probabilities accurately from limited amounts of training data. Consequently, a new method appeared to avoid problems that HMM face. However, the transition probabilities are estimated using a decision tree. Based on this method a language independent POS tagger (called TreeTagger) has been implemented. The main purpose of this work is to create the language model to adapt TreeTagger for Arabic POS tagging and lemmatization. Furthermore, different configurations have been done, namely, collecting lexical resources, as well as the annotated training corpora. In addition, we used the proposed universal tagset that consists of common POS categories of 22 different languages including Arabic. We highlight the use of this tagger via various experiments on vowelled and unvowelled text from both Modern Standard Arabic and Classical Arabic. In fact, the obtained accuracies rates are 99.4%, 92.6% and 81.9% for respectively the Quranic vowelled corpus "Al-Mus'haf", the unvowelled "Al-Mus'haf1" corpus and for the NEMLAR corpus.

Research topics

  • Natural Language Processing Techniques
  • Topic Modeling
  • Speech Recognition and Synthesis

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/it4od.2016.7479306

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.