MARATTO

article · Information

Towards Reliable Healthcare LLM Agents: A Case Study for Pilgrims during Hajj

In plain language

Large gatherings such as the Hajj bring together pilgrims from diverse linguistic and cultural backgrounds who frequently face challenges in obtaining medical advice. To improve the reliability of conversational artificial intelligence in critical healthcare contexts, a specialised methodology was developed. The framework uses domain-specific fine-tuning on a large language model combined with synthetic data augmentation, introducing the HajjHealthQA dataset to provide relevant guidance. To counter misinformation, the system integrates a retrieval-augmented generation module to verify uncertain responses, which improves overall performance by five percent. A secondary agent trained on health fact-checking data also validates the medical accuracy of generated text. Evaluated through quantitative and qualitative metrics, the resulting multilingual chatbot achieves competitive outcomes against existing models and reliably addresses a broad range of pilgrim health queries.

Key takeaways

  • A dedicated dataset named HajjHealthQA was introduced alongside synthetic data augmentation to fine-tune a healthcare conversational agent.
  • Integrating a retrieval-augmented generation module to validate uncertain answers improved the model's performance by five percent.
  • A secondary artificial intelligence agent trained on health fact-checking data provides additional medical validation to reduce misinformation.
  • The system achieved competitive results against state-of-the-art models in delivering accurate and reliable guidance for pilgrim queries.

Why it matters

During mass gatherings like the Hajj, language barriers and high demand can severely restrict access to healthcare advice. Developing conversational tools that can fact-check their own medical responses helps prevent the spread of harmful misinformation, providing diverse populations with dependable, readily accessible health support during critical periods.

Commercialisation angle

This applied research demonstrates an AI system tailored for pilgrims and public health organisers facing multilingual medical inquiries during mass gatherings. The technology incorporates verification components to produce trustworthy outputs, positioning it as an applied and tested proof-of-concept that could serve as the foundation for deployable digital health assistants.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

There is a pressing need for healthcare conversational agents with domain-specific expertise to ensure the provision of accurate and reliable information tailored to specific medical contexts. Moreover, there is a notable gap in research ensuring the credibility and trustworthiness of the information provided by these healthcare agents, particularly in critical scenarios such as medical emergencies. Pilgrims come from diverse cultural and linguistic backgrounds, often facing difficulties in accessing medical advice and information. Establishing an AI-powered multilingual chatbot can bridge this gap by providing readily available medical guidance and support, contributing to the well-being and safety of pilgrims. In this paper, we present a comprehensive methodology aimed at enhancing the reliability and efficacy of healthcare conversational agents, with a specific focus on addressing the needs of Hajj pilgrims. Our approach leverages domain-specific fine-tuning techniques on a large language model, alongside synthetic data augmentation strategies, to optimize performance in delivering contextually relevant healthcare information by introducing the HajjHealthQA dataset. Additionally, we employ a retrieval-augmented generation (RAG) module as a crucial component to validate uncertain generated responses, which improves model performance by 5%. Moreover, we train a secondary AI agent on a well-known health fact-checking dataset and use it to validate medical information in the generated responses. Our approach significantly elevates the chatbot’s accuracy, demonstrating its adaptability to a wide range of pilgrim queries. We evaluate the chatbot’s performance using quantitative and qualitative metrics, highlighting its proficiency in generating accurate responses and achieve competitive results compared to state-of-the-art models, in addition to mitigating the risk of misinformation and providing users with trustworthy health information.

Research topics

  • Travel-related health issues
  • Blood donation and transfusion practices

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.3390/info15070371

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.