MARATTO

article

Comprehensive Explainable Acne Severity Assessment via Ensemble Multi-Task Deep Learning and Interactive LLM Justification

Abstract

Acne vulgaris affects nearly 85% of adolescents and young adults worldwide and is often associated with considerable psychological burden, making accurate severity assessment critical for effective treatment planning. Traditional dermatologist grading, however, suffers from subjectivity and limited inter-rater reliability (<tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$60-70 \%$</tex>), restricting its consistency in clinical and telemedicine contexts. To address these challenges, we propose an explainable multi-task deep learning framework that simultaneously classifies acne severity and quantifies lesion counts. The framework integrates ResNet50, EfficientNetB3, and DenseNet121, enhanced with Gradient-weighted Class Activation Mapping, region-of-interest analysis, and entropy-based confidence estimation. Large language model (LLM) validation further strengthens clinical alignment by verifying diagnostic relevance and interpretability. Experiments conducted on 1,461 patient images across four severity levels and externally validated on 998 patient images achieved an average accuracy of 82.6%, with DenseNet121 achieving the highest performance (85.3%). EfficientNet-B3 delivered superior lesion counting with a mean absolute error of 3.63, while confidence scores were notably robust in extreme severity cases (95.4% for clear skin, 95.8% for very severe). These results demonstrate that the proposed framework delivers clinically reliable, explainable, and scalable acne severity assessment. This study provides a pathway toward real-world adoption of AI-driven dermatology tools in precision healthcare by bridging diagnostic accuracy with explainability.

Research topics

  • Acne and Rosacea Treatments and Effects
  • Traditional Chinese Medicine Studies
  • Dermatologic Treatments and Research

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/imcom69009.2026.11360942

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.