article · Journal of Hydroinformatics
This research evaluates multiple machine learning techniques to predict the water quality index and classify water conditions using data collected from several monitoring stations. The dataset incorporated parameters including temperature, specific conductance, salinity, dissolved oxygen, depth, pH, and turbidity, which were preprocessed through imputation, feature scaling, and categorical encoding. The study tested artificial neural networks, decision trees, support vector machines, random forests, XGBoost, and long short-term memory networks. Among these algorithms, XGBoost and long short-term memory models proved superior to the alternatives. XGBoost achieved classification accuracy between 99.07 and 99.99 percent, while the long short-term memory model reached an R-squared value of 0.9999. These models demonstrated enhanced robustness, predictive precision, and generalisation capabilities relative to traditional analytical approaches when handling complex multivariate water data.
Accurate monitoring of water resources is critical for public health and environmental sustainability. By proving that advanced computational algorithms can interpret complex environmental metrics with near-perfect accuracy, this research provides environmental managers and policymakers with powerful analytical methods to better detect pollution, manage natural resources, and safeguard drinking water supplies.
The models could enable automated, scalable digital tools for environmental management and regulatory water monitoring. Potential users include environmental protection agencies, municipal utilities, and water resource managers. As the models have been trained and tested on station datasets with high predictive accuracy, the research sits at an applied, validated stage, though commercial deployment would require integration into operational monitoring software.
AI-generated from the published abstract. Always read the original work before citing.
ABSTRACT This study presents an in-depth analysis of machine learning (ML) techniques for predicting water quality index and water quality classification using a dataset containing water quality metrics such as temperature, specific conductance, salinity, dissolved oxygen, depth, pH, and turbidity from multiple monitoring stations. Data preprocessing included imputation for missing values, feature scaling, and categorical encoding, ensuring balanced input features. This research evaluated artificial neural networks, decision trees, support vector machines, random forests, XGBoost, and long short-term memory (LSTM) networks. Results demonstrate that XGBoost and LSTM significantly outperformed other models, with XGBoost achieving an accuracy range of 99.07–99.99% and LSTM attaining an R2 of 0.9999. Compared with prior studies, our approach enhances predictive accuracy and robustness, showcasing advanced generalization capabilities. The proposed models exhibit significant improvements over traditional methods in handling complex, multivariate water quality data, positioning them as promising tools for water quality prediction and environmental management. These findings underscore the potential of ML for developing reliable, scalable water quality monitoring solutions, providing valuable insights for policymakers and environmental managers dedicated to sustainable water resource management.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.2166/hydro.2025.290
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.