article
Automatic recognition of handwritten text in images remains a major challenge in computer vision, due to the variability of handwriting styles, visual noise, and complex page layouts. This paper introduces a hybrid approach that combines YOLOv8 for detecting text regions with TrOCR, a pretrained Transformer-based model, for text recognition. Our method is applied to the RIMES 2011-line dataset, which contains lines of handwritten French text. The first stage involves accurately locating text lines using YOLOv8, followed by extraction and transcription of each line with TrOCR. Experimental results show that this combination achieves a mAP@0.5 of 98.7% for detection, while also revealing certain limitations in the linguistic quality of the transcriptions produced. The proposed approach demonstrates the effectiveness of a sequential detection-recognition architecture, while also highlighting open challenges in contextual transcription.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/acdsa67686.2026.11468250
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.