article · Zenodo (CERN European Organization for Nuclear Research)
The rapid growth of the dark web has transformed it into a major platform for the exchange of stolen data, malware, exploits, and other illicit cyber services. However, the unstructured, slang-driven, and highly contextual nature of dark web communications presents significant challenges for conventional cyber threat detection techniques. This study proposes a transformer-based Natural Language Processing (NLP) framework for extracting and interpreting emerging cyber threats from dark web forums and marketplaces. A large-scale dataset comprising over four million forum posts collected from seven prominent dark web platforms was preprocessed and analysed to identify discussions related to fraud, malware, data leaks, exploits, drugs, weapons, and other illicit activities. Two transformer-based models were developed to perform forum classification and threat-type detection. Experimental results show that the BERT model achieved 78.42% accuracy in forum classification, while the DistilBERT model achieved 97.81% accuracy with an average F1-score of 0.88 for threat-type detection. The proposed framework demonstrates the effectiveness of transformer-based NLP models in extracting actionable cyber threat intelligence from unstructured dark web communications and provides a scalable foundation for proactive cyber threat monitoring and early warning systems.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.21818771
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.