article · American Journal of Interdisciplinary Research and Innovation
Internet platforms face continuous threats from malicious URLs used for phishing, malware distribution, and broader cybercrime, endangering both individuals and organisations. Machine learning models offer a way to identify these dangerous links. Using a public dataset of over 651,000 labelled URLs, data was prepared using stratified sampling, label encoding, and Term Frequency-Inverse Document Frequency vectorisation. Five algorithms were evaluated, namely Support Vector Machine, K-Nearest Neighbours, Naïve Bayes, Random Forest, and Extreme Gradient Boosting. Evaluation based on accuracy, precision, recall, and F1 score revealed that the Random Forest model achieved the highest classification accuracy at 95 per cent. Furthermore, a web-based detection system was constructed to demonstrate how these models perform in real-time cybersecurity settings, showing that ensemble learning methods provide an effective and dependable mechanism for malicious URL identification.
Malicious web links are primary vectors for phishing scams, malware infections, and digital fraud, threatening online safety for individuals and enterprises alike. Automated detection tools that evaluate web addresses with high accuracy can identify threats before users open compromised sites, providing an essential layer of defence for digital communication, online commerce, and everyday internet use.
The technology enables automated, real-time filtering of malicious links for web browsers, corporate security networks, and messaging services. Prospective users include cybersecurity vendors, enterprise IT departments, and everyday internet users needing defence against phishing. Having been integrated and demonstrated within a functional web-based detection system, the work represents an applied and tested technology positioned close to operational deployment.
AI-generated from the published abstract. Always read the original work before citing.
The internet has rapidly evolved in its communication, commerce and information sharing making it a huge platform for cyber threats, particularly malicious URLs. They pose a serious threat to individuals and to organisations. Phishing attacks, malware distribution and other types of cybercrime frequently are carried out through malicious URLs. In this research, we have created and tested the machine learning models to detect malicious URLs. The labeled URLs used were obtained from a public dataset with more than 651,000 labeled URLs. The dataset was prepared for classification by applying data pre-processing techniques like stratified sampling, label encoding and Term Frequency–Inverse Document Frequency (TF-IDF) vectorization. To train and test the algorithms, five machine learning were used: Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Naïve Bayes (NB), Random Forest (RF) and Extreme Gradient Boosting (XGBoost), which were trained and evaluated by the metrics of accuracy, precision, recall and F1 score. The results indicated that the Random Forest model had the highest classification accuracy (95%) as compared to the other models. Moreover, a web based malicious URL detection system was developed to demonstrate the actual application of the developed models in real time cyber security scenarios. The study provides a conclusion that the machine learning techniques, particularly ensemble learning techniques can be considered an effective and reliable technique to detect malicious URL.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.54536/ajiri.v5i3.8109
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.