article · Scientific Reports
Large language models have potential utility in medical education, particularly in assisting students as they prepare for licensing examinations. A study evaluated seven publicly accessible models across 423 simulated UK medical board exam questions from nine examinations, spanning subjects such as surgery and paediatrics. The evaluated systems included ChatGPT-3.5, ChatGPT-4, Bard, Perplexity, Claude, Bing, and Claude Instant. Performance varied significantly across the tools, with ChatGPT-4 achieving the highest score at 78.2 per cent, followed by Bing at 67.2 per cent, Claude at 64.4 per cent, and Perplexity scoring the lowest at 56.1 per cent. Across all models, accuracy was higher on multiple-choice questions than on true-false or choose-n formats. While these tools show promise for training, observed errors demonstrate clear limitations, highlighting that further refinement and research are necessary before relying on them heavily in medical education curricula.
Medical trainees require accurate, dependable tools to support their learning. By benchmarking existing public artificial intelligence models against standard UK medical board examinations, this research demonstrates both the educational promise and the current accuracy limits of these systems, guiding safe adoption in professional training environments.
The findings inform developers and educational institutions aiming to build artificial intelligence study aids for medical students. Public models demonstrate applied testing on simulated exams, but remaining errors indicate that current tools are not ready for unsupervised deployment. Real-world implementation will require further refinement, potentially through specialty-specific systems.
AI-generated from the published abstract. Always read the original work before citing.
Large language models (LLMs) like ChatGPT have potential applications in medical education such as helping students study for their licensing exams by discussing unclear questions with them. However, they require evaluation on these complex tasks. The purpose of this study was to evaluate how well publicly accessible LLMs performed on simulated UK medical board exam questions. 423 board-style questions from 9 UK exams (MRCS, MRCP, etc.) were answered by seven LLMs (ChatGPT-3.5, ChatGPT-4, Bard, Perplexity, Claude, Bing, Claude Instant). There were 406 multiple-choice, 13 true/false, and 4 "choose N" questions covering topics in surgery, pediatrics, and other disciplines. The accuracy of the output was graded. Statistics were used to analyze differences among LLMs. Leaked questions were excluded from the primary analysis. ChatGPT 4.0 scored (78.2%), Bing (67.2%), Claude (64.4%), and Claude Instant (62.9%). Perplexity scored the lowest (56.1%). Scores differed significantly between LLMs overall (p < 0.001) and in pairwise comparisons. All LLMs scored higher on multiple-choice vs true/false or "choose N" questions. LLMs demonstrated limitations in answering certain questions, indicating refinements needed before primary reliance in medical education. However, their expanding capabilities suggest a potential to improve training if thoughtfully implemented. Further research should explore specialty specific LLMs and optimal integration into medical curricula.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1038/s41598-024-68996-2
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.