MARATTO

article · Oral Diseases

Diagnostic Performance of ChatGPT‐4o and DeepSeek‐3 Differential Diagnosis of Complex Oral Lesions: A Multimodal Imaging and Case Difficulty Analysis

202522 citationsSuez University

In plain language

Artificial intelligence models show potential in diagnostic medicine, yet their efficacy in identifying complex oral lesions requires clear evaluation. A comparative assessment tested the diagnostic accuracy of ChatGPT-4o, DeepSeek-3, and four board-certified oral medicine specialists using 80 standardised clinical vignettes featuring clinical images and radiographs. Specialist clinicians consistently achieved the highest diagnostic accuracy across all evaluations. Between the models, DeepSeek-3 significantly outperformed ChatGPT-4o at the Top-3 differential diagnosis level, demonstrating stronger performance in challenging and inflammatory cases despite operating as a text-only system. While multimodal imaging improved overall diagnostic outcomes, increasing case difficulty reduced Top-1 accuracy across tools. Text-based language models can provide effective diagnostic reasoning with lower hallucination rates, but expert clinical oversight remains essential when evaluating complex oral conditions.

Key takeaways

  • Board-certified oral medicine specialists consistently outperformed both artificial intelligence models in diagnostic accuracy.
  • DeepSeek-3 significantly exceeded ChatGPT-4o in Top-3 diagnostic accuracy and showed greater resilience in high-difficulty and inflammatory cases.
  • Multimodal imaging improved diagnostic accuracy, whereas higher case difficulty reduced Top-1 diagnostic performance.
  • The text-only DeepSeek-3 model demonstrated stronger structured reasoning and fewer hallucinations than the multimodal ChatGPT-4o model.
  • Expert medical oversight remains necessary when deploying language models for complex diagnostic support.

Why it matters

Accurate oral disease diagnosis often depends on specialised clinical expertise that is not always readily accessible. These findings demonstrate that advanced text-based artificial intelligence can assist in identifying complex conditions, outperforming certain multimodal tools. However, because specialist doctors still perform best, such tools are most appropriately positioned as supportive aids rather than independent decision-makers in clinical patient care.

Commercialisation angle

The findings highlight an application pathway for text-based artificial intelligence as clinical diagnostic decision-support software for dental and oral medicine practitioners. The evaluation relies on eighty standardised retrospective vignettes rather than live clinical workflows, indicating an early-stage research level of readiness. Real-world commercial deployment would require formal integration with clinical systems and continuous supervision by qualified healthcare professionals to manage complex or high-difficulty diagnostic scenarios safely.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

BACKGROUND: AI models like ChatGPT-4o and DeepSeek-3 show diagnostic promise, but their reliability in complex, image-based oral lesions remains unclear. This study aimed to evaluate and compare the diagnostic accuracy of ChatGPT-4o and DeepSeek-3 despite their differing modalities against oral medicine (OM) experts across varied lesion types and case difficulty levels. METHODS: Eighty standardized clinical vignettes derived from real-world oral disease cases, including clinical images/radiographs, were evaluated. Differential diagnoses were generated by ChatGPT-4o, DeepSeek-3, and four board-certified OM specialists, with accuracy assessed at Top-1, Top-3, and Top-5 levels. RESULTS: OM specialists consistently achieved the highest diagnostic accuracy. However, DeepSeek-3 significantly outperformed ChatGPT-4o at the Top-3 level (p = 0.0153) and showed greater robustness in high-difficulty and inflammatory cases despite its text-only modality. Multimodal imaging enhanced diagnostic accuracy. Regression analysis indicated lesion type and imaging modality as positive predictors, while diagnostic difficulty negatively impacted Top-1 performance. CONCLUSIONS: Remarkably, the text-only DeepSeek-3 model exceeded the diagnostic performance of the multimodal ChatGPT-4o model for complex oral lesions, highlighting its structured reasoning capabilities and reduced hallucination rate. These findings underscore the potential of non-vision LLMs in diagnostic support, emphasizing the critical need for expert oversight in complex scenarios.

Research topics

  • Artificial Intelligence in Healthcare and Education
  • Radiomics and Machine Learning in Medical Imaging
  • Clinical Reasoning and Diagnostic Skills

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1111/odi.70007

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.