MARATTO

article · Diagnostics

Polycystic Ovary Syndrome Detection Machine Learning Model Based on Optimized Feature Selection and Explainable Artificial Intelligence

2023104 citationsOpen accessSuez University

In plain language

Polycystic ovary syndrome is a common and severe condition affecting women globally, linked to increased risks of gestational and type 2 diabetes. Detecting the condition early helps healthcare systems reduce these long-term complications. Researchers evaluated several machine learning models, including logistic regression, random forest, decision trees, support vector machines, and gradient boosting algorithms, alongside feature selection techniques and Bayesian optimisation. To resolve dataset imbalance, synthetic minority oversampling was combined with edited nearest neighbour techniques. The research also incorporated local and global explainable artificial intelligence methods to improve transparency and trust in automated diagnosis. Tested on a benchmark polycystic ovary syndrome dataset using two split ratios, an ensemble stacking model combining top-performing base algorithms with recursive feature elimination achieved an accuracy of 100 percent.

Key takeaways

  • A stacking ensemble machine learning model achieved 100 percent diagnostic accuracy on a benchmark polycystic ovary syndrome dataset.
  • A combination of oversampling and nearest neighbour filtering was applied to address class imbalance in the training data.
  • Bayesian optimisation and recursive feature elimination were employed to identify the best-performing model parameters and features.
  • Local and global explainable artificial intelligence methods were integrated to ensure diagnostic decisions remain interpretable and trustworthy.

Why it matters

Polycystic ovary syndrome is a widespread health issue that frequently leads to severe metabolic complications when left unmanaged. Establishing accurate, interpretable computational diagnostic tools can support healthcare providers in identifying the condition at an earlier stage, allowing for timely clinical intervention and improved patient care.

Commercialisation angle

The work could inform the development of clinical decision-support software for healthcare providers screening patients for polycystic ovary syndrome. By providing explainable outputs alongside automated classifications, such tools could assist medical professionals during diagnostics. However, the evaluation relies solely on a single benchmark dataset, indicating that the technology is at an applied research stage and requires clinical validation before real-world adoption.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Polycystic ovary syndrome (PCOS) has been classified as a severe health problem common among women globally. Early detection and treatment of PCOS reduce the possibility of long-term complications, such as increasing the chances of developing type 2 diabetes and gestational diabetes. Therefore, effective and early PCOS diagnosis will help the healthcare systems to reduce the disease's problems and complications. Machine learning (ML) and ensemble learning have recently shown promising results in medical diagnostics. The main goal of our research is to provide model explanations to ensure efficiency, effectiveness, and trust in the developed model through local and global explanations. Feature selection methods with different types of ML models (logistic regression (LR), random forest (RF), decision tree (DT), naive Bayes (NB), support vector machine (SVM), k-nearest neighbor (KNN), xgboost, and Adaboost algorithm to get optimal feature selection and best model. Stacking ML models that combine the best base ML models with meta-learner are proposed to improve performance. Bayesian optimization is used to optimize ML models. Combining SMOTE (Synthetic Minority Oversampling Techniques) and ENN (Edited Nearest Neighbour) solves the class imbalance. The experimental results were made using a benchmark PCOS dataset with two ratios splitting 70:30 and 80:20. The result showed that the Stacking ML with REF feature selection recorded the highest accuracy at 100 compared to other models.

Research topics

  • Ovarian function and disorders

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.3390/diagnostics13081506

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.