MARATTO

article

Voice Pathology Detection Using Convolutional Neural Network and Data Augmentation

Abstract

Detecting and identifying voice disorders, which frequently result in diminished vocal quality, is essential for effective treatment and improved patient outcomes. This paper focuses on developing a Deep Learning-based approach to identify voice pathologies. The proposed approach combines the use of Convolutional Neural Network (CNN) with data augmentation techniques. Mel Frequency Cepstral Coefficients (MFCCs) are utilized as input features to the CNN model to perform binary classification of voices (normal or pathological). Experiments were conducted using two types of voice recordings from the Saarbrücken Voice Database (SVD), including vowels and sentences. The evaluation results demonstrate the superiority of our approach in in terms of accuracy for pathological voice detection compared to SVM appraoch and KNN techniques.

Research topics

  • Voice and Speech Disorders
  • Respiratory and Cough-Related Research
  • Phonocardiography and Auscultation Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/ic_aset65966.2025.11231992

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.