article
Medical Image Captioning is an emerging field in clinical AI aimed at generating structured, explainable textual reports from radiological images to support automated diagnostics and decision-making. This study presents a domain-specific model that fine-tunes the Bootstrapped Language-Image Pretraining (BLIP) framework using a subset of the Radiology Objects in COntext (ROCO) dataset annotated with Unified Medical Language System (UMLS) concepts. The proposed pipeline incorporates comprehensive image and caption preprocessing, a Vision Transformer (ViT) encoder-decoder architecture, and hyperparameter optimization through Particle Swarm Optimization (PSO) to enhance convergence and generalization. Evaluation using BLEU, METEOR, and ROUGE-L metrics revealed notable improvements after optimization, with BLEU-1 reaching 0.8100, BLEU-4 reaching 0.7700, METEOR increasing to 0.6300, and ROUGE-L to 0.8550. These results underscore the effectiveness of integrating PSO with BLIP for generating clinically accurate and semantically rich radiology reports.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/comnet68251.2025.11325450
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.