review · Informatica
The problem is made more difficult by the fact that the recognition of ancient handwritten Arabic script (AHR) is written in cursive, has different historical styles, and the manuscripts are often damaged. In addition, ancient handwriting does not follow modern standards of handwriting spacing, which includes overly spaced-out words and overly complex diacritics, which makes it extremely difficult to process. This irregularity causes ambiguity in character segmentation and word boundaries, increasing the error rate in automatic recognition systems. Even with modern advancements in deep learning through the use of CNNs, LSTMs, and hybrid models, AHR is still extremely complex and requires a lot of exploration. Some recent models have achieved accuracy between 70% and 90% on modern Arabic datasets, but performance drops to 50%–75% when applied to ancient texts due to noise, script variation, and limited annotated data. The article consolidates the major issues and recent developments with regard to dataset constraints, preprocessing requirements, and machine learning methodologies. This review is based on the analysis of over 50 peer-reviewed papers published between 2016 and 2024. It is also focused on the importance of deep learning in the image feature extraction by CNNs, sequential feature modeling by LSTMs, and combination of both – hybrids. For instance, CNN-LSTM architectures have shown promising results on historical scripts with limited training data. With so little annotated data available, it concentrates on the augmentation of datasets and creation of synthetic data. Techniques such as elastic distortions, GAN-generated samples, and noise injection are discussed as potential solutions. This work aims to improve the accuracy and scalability of AHR through analysis of existing techniques and identification of the gaps for further research to aid in digitization and analysis of manuscripts to safeguard them as a part of cultural heritage. In particular, this review highlights the lack of standardized benchmarks and the need for multilingual ancient Arabic datasets to support reproducible research.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.31449/inf.v49i28.8920
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.