article
Accurate diagnosis of oral lesions remains challenging due to overlapping clinical features and the high risk of missing malignancies. Multimodal large language models (LLMs), such as ChatGPT-4 and Google's Gemini Pro 2.5, can assist clinicians by integrating textual and visual data. This study compared these models with human experts in diagnosing oral lesions and quantified the added value of clinical images (photographs radiographs) on diagnostic accuracy. A total of 160 case vignettes with intraoral images were evaluated using ChatGPT-4 and Gemini Pro 2.5, with Top-1, Top-3, and Top-5 accuracy metrics benchmarked against two oral medicine specialists. Each model was tested with and without images, and analyses included Cochran's Q, McNemar tests with Bonferroni correction, Cohen's <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$h$</tex>, and logistic regression. With images, ChatGPT-4 achieved 63.7 % Top-1 accuracy versus Gemini's 71.2 % and experts' 87.5 %. ChatGPT-4 improved significantly with image input (+13.8 points, <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$p=0.017$</tex>), while Gemini's gain was smaller and non-significant. Both reached <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\sim 95 \%$</tex> Top-3 accuracy, closing the gap with experts. Visual input was most beneficial in high-difficulty and morphologically complex cases, while radiographs offered limited additional value. These findings underscore the promise of multimodal LLMs as assistive tools in oral diagnostics and the need for cautious, evidence-based clinical integration.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/icicis66182.2025.11313191
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.