review
<title>Abstract</title> Large Language Models (LLMs) have shown potential to transform medical specialties but there are uncertainties in their performance. This systematic review aims to evaluate LLMs utilization, impacts, and challenges across 19 medical specialties. We systematically searched on Scopus and Web of Science for journal articles employing LLMs for medical specialties. 5,790 peer-reviewed studies were identified, of which 84 were included in this systematic review. The results revealed the most used evaluation metric for assessing LLM performance is the accuracy metric, followed by the F1-score. The LLM scores for the evaluation 1 metrics varied based on the task of utilizing the LLM, where five main applications were identified. The main positive impact reported was enhancing medical processes efficiency, while the most reported negative impact was inconsistent reliability. More rigorous validation standards that consider the unique challenges of LLMs research could enhance future studies. Continued research and collaboration between computer science and medical communities will be crucial for advancing and ensuring the reliable and trustworthy application of LLMs in healthcare.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.21203/rs.3.rs-5128451/v1
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.