article · BioFactors
Early detection of breast cancer is vital for reducing mortality and improving treatment outcomes, yet acquiring balanced clinical datasets remains a persistent challenge for diagnostic tools. To address this imbalance, a framework combining random augmentation and SHapley Additive exPlanations, known as SHAP, was applied to the Wisconsin breast cancer dataset. Random augmentation was used to synthetically generate portions of the data, while SHAP evaluated the quality of specific attributes chosen to train diagnostic models. Evaluating this integrated approach across six distinct machine learning algorithms showed general performance gains of more than 3% for most models. Beyond initial diagnosis, the explainability provided by SHAP was used to inform ongoing patient care strategies, helping healthcare providers concentrate on timely and high quality disease management to lower the risk of recurrence and related medical complications.
Imbalanced medical data often hampers the reliability of diagnostic algorithms. By pairing synthetic data generation with model interpretability, diagnostic tools can achieve greater predictive accuracy. Furthermore, using explainable features helps clinicians understand model decisions, enabling more targeted and timely post diagnosis patient management to mitigate complications.
This work represents early stage applied research focused on algorithmic development using the Wisconsin breast cancer benchmark dataset. It could inform diagnostic software tools and decision support systems used by clinical oncology teams. However, the abstract does not indicate testing in live clinical workflows, indicating that real world deployment remains at an early developmental distance.
AI-generated from the published abstract. Always read the original work before citing.
Recent research indicates that early detection of breast cancer (BC) is critical in achieving favorable treatment outcomes and reducing the mortality rate associated with it. With the difficulty in obtaining a balanced dataset that is primarily sourced for the diagnosis of the disease, many researchers have relied on data augmentation techniques, thereby having varying datasets with varying quality and results. The dataset we focused on in this study is crafted from SHapley Additive exPlanations (SHAP)-augmentation and random augmentation (RA) approaches to dealing with imbalanced data. This was carried out on the Wisconsin BC dataset and the effectiveness of this approach to the diagnosis of BC was checked using six machine-learning algorithms. RA synthetically generated some parts of the dataset while SHAP helped in assessing the quality of the attributes, which were selected and used for the training of the models. The result from our analysis shows that the performance of the models used generally increased to more than 3% for most of the models using the dataset obtained by the integration of SHAP and RA. Additionally, after diagnosis, it is important to focus on providing quality care to ensure the best possible outcomes for patients. The need for proper management of the disease state is crucial so as to reduce the recurrence of the disease and other associated complications. Thus the interpretability provided by SHAP enlightens the management strategies in this study focusing on the quality of care given to the patient and how timely the care is.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1002/biof.1995
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.