MARATTO

article · DOAJ (DOAJ: Directory of Open Access Journals)

A Framework for Optimising Phishing Websites Detection via Ensemble Learning and Synthetic Minority Oversampling Techniques

2026Open accessUniversity of Uyo

Abstract

Phishing websites remain one of the most prevalent forms of cyber-attacks. They often exploit users, through deceptive web interfaces, to steal sensitive information such as login credentials, financial data, and personal records. The increasing sophistication of phishing strategies has reduced the effectiveness of traditional rule-based detection systems. This necessitates the adoption of intelligent and adaptive machine learning (ML) approaches. This study presents a robust ML-based framework for phishing websites detection using ensemble learning and Synthetic Minority Oversampling Technique (SMOTE). The proposed framework integrates comprehensive data preprocessing techniques, including constant and duplicate feature removal, Standard Scaler (Z-score normalization), and class balancing through SMOTE to improve model robustness and predictive capability. Three supervised learning algorithms, namely Decision Tree (DT), Random Forest (RF), and Extreme Gradient Boosting (XGBoost), were implemented and evaluated using an 8:2 train-test split, 5-fold cross-validation, and GridSearchCV-based hyperparameter optimization. Experimental results demonstrate that ensemble models outperform the standalone DT model across all evaluation metrics. Among the evaluated models, RF achieved the best overall performance with an accuracy of 97.13%, recall of 96.07%, F1-score of 95.86%, and area under the curve-receiver operating characteristics (AUC-ROC) of 0.9951 on the baseline dataset. Following the application of SMOTE, recall improved further to 96.46%, indicating enhanced capability in identifying phishing websites. Additionally, SHapley Additive exPlanations (SHAP) was incorporated to improve model interpretability and identify the most influential phishing indicators. The findings reveal that domain age and URL structural characteristics significantly influence phishing detection. Overall, the findings demonstrate the effectiveness and interpretability of the evaluated tree-based framework within the structured-feature dataset considered in this study. However, the age of the dataset limits the extent to which the findings can be generalized to contemporary phishing environments. Future work should therefore evaluate the framework on recent and continuously evolving datasets and investigate its robustness against emerging phishing strategies and changes in website characteristics.

Research topics

  • Spam and Phishing Detection
  • Cybercrime and Law Enforcement Studies
  • Advanced Malware Detection Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.22102917

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.