MARATTO

article · Scientific Reports

AI chatbots show promise but limitations on UK medical exam questions: a comparative performance study

In plain language

Large language models have potential utility in medical education, particularly in assisting students as they prepare for licensing examinations. A study evaluated seven publicly accessible models across 423 simulated UK medical board exam questions from nine examinations, spanning subjects such as surgery and paediatrics. The evaluated systems included ChatGPT-3.5, ChatGPT-4, Bard, Perplexity, Claude, Bing, and Claude Instant. Performance varied significantly across the tools, with ChatGPT-4 achieving the highest score at 78.2 per cent, followed by Bing at 67.2 per cent, Claude at 64.4 per cent, and Perplexity scoring the lowest at 56.1 per cent. Across all models, accuracy was higher on multiple-choice questions than on true-false or choose-n formats. While these tools show promise for training, observed errors demonstrate clear limitations, highlighting that further refinement and research are necessary before relying on them heavily in medical education curricula.

Key takeaways

  • ChatGPT-4 achieved the highest accuracy on UK medical exam questions at 78.2 per cent, while Perplexity scored lowest at 56.1 per cent.
  • Overall performance differed significantly across the seven evaluated large language models.
  • All tested models achieved higher accuracy on multiple-choice questions compared to true-false or choose-n formats.
  • Current limitations show that models require further refinement before they can be primarily relied upon in medical education.

Why it matters

Medical trainees require accurate, dependable tools to support their learning. By benchmarking existing public artificial intelligence models against standard UK medical board examinations, this research demonstrates both the educational promise and the current accuracy limits of these systems, guiding safe adoption in professional training environments.

Commercialisation angle

The findings inform developers and educational institutions aiming to build artificial intelligence study aids for medical students. Public models demonstrate applied testing on simulated exams, but remaining errors indicate that current tools are not ready for unsupervised deployment. Real-world implementation will require further refinement, potentially through specialty-specific systems.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Large language models (LLMs) like ChatGPT have potential applications in medical education such as helping students study for their licensing exams by discussing unclear questions with them. However, they require evaluation on these complex tasks. The purpose of this study was to evaluate how well publicly accessible LLMs performed on simulated UK medical board exam questions. 423 board-style questions from 9 UK exams (MRCS, MRCP, etc.) were answered by seven LLMs (ChatGPT-3.5, ChatGPT-4, Bard, Perplexity, Claude, Bing, Claude Instant). There were 406 multiple-choice, 13 true/false, and 4 "choose N" questions covering topics in surgery, pediatrics, and other disciplines. The accuracy of the output was graded. Statistics were used to analyze differences among LLMs. Leaked questions were excluded from the primary analysis. ChatGPT 4.0 scored (78.2%), Bing (67.2%), Claude (64.4%), and Claude Instant (62.9%). Perplexity scored the lowest (56.1%). Scores differed significantly between LLMs overall (p < 0.001) and in pairwise comparisons. All LLMs scored higher on multiple-choice vs true/false or "choose N" questions. LLMs demonstrated limitations in answering certain questions, indicating refinements needed before primary reliance in medical education. However, their expanding capabilities suggest a potential to improve training if thoughtfully implemented. Further research should explore specialty specific LLMs and optimal integration into medical curricula.

Research topics

  • Artificial Intelligence in Healthcare and Education
  • COVID-19 diagnosis using AI
  • Radiomics and Machine Learning in Medical Imaging

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1038/s41598-024-68996-2

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.