article
Optical Character Recognition helps to easily conversion of typed, printed, or handwritten text images into editable and searchable data. Working on optical Character recognition is challenging due to the variety of writing scripts, the presence of visually similar characters, and the degradation of documents. In other natural language such as Latin, Arabic, Chinese, and Amharic there are number of works, but different dialects Guragegna language has received little attention. In this research we proposed new model that recognize multi-dialect Guragegna printed documents. Although Guragegna belongs to the same language family as Amharic, it has unique features, characters that are not recognized by existing optical character recognition models designed for Amharic. The proposed model includes four components including prepossessing, segmentation, normalization, and recognition. The dataset has been collected from various sources, using a digital camera. These images undergo preprocessing before being fed into novel hybrid deep learning model. We conducted experiments and compare other three models, including convolutional neural networks and recurrent neural networks and a hybrid model. Based on the experimental accuracy results, our newly proposed hybrid model outperforms with the accuracy of 93.19%.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/ict4da62874.2024.10777113
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.