MARATTO

article

Enhancing Credit Risk Assessment with Machine Learning Algorithms: A Comparative Study of Label Encoding and One-Hot Encoding Across Multiple Classifiers

Abstract

This study focuses on using machine learning for credit risk assessment, which is crucial in making financial decisions. As financial systems become more complex, traditional methods are no longer enough to accurately assess risk. Machine learning models offer the ability to uncover hidden patterns and relationships within the data, providing a more reliable and efficient way to assess credit risk. In our research, we compare the performance of several machine learning models to understand how they perform with different data preprocessing techniques. Specifically, we compare the effects of Label Encoding and One-Hot Encoding on these models. The study highlights the importance of preprocessing, especially when working with imbalanced data, which is common in credit risk assessment. XGBoost yields the best results, with CatBoost also showing strong performance. By examining these models and encoding methods, this research offers valuable insights into improving predictive accuracy and ensuring more reliable credit risk evaluations.

Research topics

  • Financial Distress and Bankruptcy Prediction
  • Imbalanced Data Classification Techniques
  • Credit Risk and Financial Regulations

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/niss66502.2025.00030

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.