article · Procedia Computer Science
The application of deep learning to character recognition has significantly advanced text digitization. However, challenges remain for less common languages like Lingala, widely spoken in the Democratic Republic of Congo and the Republic of Congo. This paper introduces an innovative integration of LayoutLM, which interprets textual content within its spatial layout, and Faster R-CNN, which accurately detects text areas in images. This combination effectively converts scanned documents into editable digital text. We compiled a comprehensive Lingala text dataset from various sources to train and validate our model. This approach enhances the accurate interpretation of Lingala documents and promotes broader access to digital information in underrepresented languages. The project encourages further community involvement to improve and expand our model’s capabilities.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1016/j.procs.2025.03.017
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.