MARATTO

article · Scientific Reports

Prompt-dependent performance of multimodal AI model in oral diagnosis: a comprehensive analysis of accuracy, narrative quality, calibration, and latency versus human experts

202517 citationsOpen accessSinai University

In plain language

This study evaluated how different prompting strategies affect the performance of a multimodal artificial intelligence model, Gemini Pro 2.5, in diagnosing oral lesions. Researchers tested 300 histopathology-verified clinical cases using three prompt styles: Direct, Chain-of-Thought, and Self-Reflection, comparing results against board-certified oral medicine specialists. Human specialists achieved the highest top single diagnosis accuracy at 61 percent, and the model could not match human performance on low-difficulty cases. However, Chain-of-Thought prompting yielded the best artificial intelligence performance, achieving 82 percent accuracy across the top three differential diagnoses, along with the highest narrative explanation quality and superior probability calibration. While direct prompting operated with the lowest computational latency, structured reasoning prompts delivered better diagnostic reliability and richer clinical context. The results demonstrate that prompt structure substantially influences the diagnostic output and interpretability of multimodal models in oral medicine.

Key takeaways

  • Human oral medicine experts achieved the highest top single diagnosis accuracy at 61 percent, outperforming the model on low-difficulty cases.
  • Chain-of-Thought prompting delivered the best model performance, reaching 82 percent top-three diagnostic accuracy and the highest narrative quality.
  • Chain-of-Thought reasoning also provided the best probability calibration among the tested artificial intelligence prompts.
  • Direct prompting was the fastest approach computationally, but structured prompting offered better diagnostic recall and explanation depth.

Why it matters

Using artificial intelligence to assist medical diagnoses requires outputs that clinicians can trust and interpret. By showing that prompt design directly alters diagnostic accuracy, certainty, and explanation quality, this work highlights how structured reasoning methods can make multimodal models safer and more effective clinical partners, particularly for evaluating complex oral conditions.

Commercialisation angle

This research provides applied insights for software developers building diagnostic decision-support tools for dental and oral medicine clinicians. By demonstrating that Chain-of-Thought prompting improves top-three diagnostic accuracy and explanation quality, the findings inform prompt engineering strategies within clinical software. As an applied accuracy study tested against verified cases and human experts, it sits at an intermediate validation stage prior to real-world clinical workflow integration.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Prompt design is a critical yet underexplored factor influencing the diagnostic performance of large language models (LLMs). Gemini Pro 2.5 shows promise in multimodal reasoning, but no prior study has systematically compared prompt structures in oral datasets against expert benchmarks. This study aimed to evaluate the diagnostic performance of a multimodal LLM (Gemini Pro 2.5) under different prompting strategies compared with oral medicine experts using prospective, histopathology-verified clinical vignettes. In a prospective, paired diagnostic accuracy study, Gemini pro 2.5 (a multimodal LLM) was evaluated under three prompting strategies: Direct (P-1), Chain-of-Thought (P-2), and Self-Reflection (P-3) on 300 oral lesion cases with histopathologic confirmation. Each prompt was applied to identical inputs and compared against diagnoses from board-certified oral medicine specialists. Accuracy, rubric-based narrative quality, probability calibration, and computational efficiency were assessed under STARD-AI guidelines. Human experts achieved the highest Top-1 accuracy (61%), but Chain-of-Thought prompting (P-2) led AI performance in Top-3 accuracy (82%) and produced the highest explanation quality (mean rubric score 8.49/10). No AI prompt matched human performance in low-difficulty cases. P-2 also showed the best calibration (Brier score 0.238) compared to P-1 and P-3. Resource-wise, Direct prompting was fastest, but longer outputs modestly improved Top-3 recall. Mixed-effects modeling confirmed that AI performance varied significantly by prompt structure, highlighting context-specific trade-offs. Prompt structure significantly affects the diagnostic performance and interpretability of AI-generated differentials in oral lesion diagnosis. While expert clinicians remain superior in straightforward cases, structured prompting, particularly Chain-of-Thought, may enhance AI reliability in complex diagnostic scenarios. These findings support the integration of prompt engineering into AI-assisted diagnostic tools to augment clinical decision-making in oral medicine.

Research topics

  • Artificial Intelligence in Healthcare and Education
  • Radiomics and Machine Learning in Medical Imaging
  • Clinical Reasoning and Diagnostic Skills

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1038/s41598-025-22979-z

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.