MARATTO

article · International Journal of Data Science and Analytics

Applying metaheuristic optimization for insider-threat detection using natural language processing

Abstract

Maintaining trust among employees, employers, and institutions is fundamental to business and research, yet the rise of always-online digital systems and expanding workforces poses new risks to operational integrity. Fluctuating work environments, evolving motivations, and gaps in training can leave organizations vulnerable to insider threats, including inadvertent data leaks and intentional exfiltration. In this study, an investigation of natural language processing (NLP) methods applied to HTTP activity logs to identify potential insider threats through behavioral patterns is conducted. Two experiments were conducted using TF-IDF and Word2Vec text representations, each combined with an XGBoost classifier whose hyperparameters were optimized using a newly proposed iteration stagnation-aware variable neighborhood search (ISAVNS) metaheuristic. The ISAVNS introduces a stagnation-detection mechanism that enables adaptive recovery during optimization, improving exploration and convergence stability. Evaluation on publicly available insider-threat datasets confirmed the high effectiveness of the proposed framework, with the TF-IDF-based model reaching an accuracy of 97.63% and the Word2Vec-based counterpart attaining 97.71%.

Research topics

  • Software System Performance and Reliability
  • Mental Health via Writing
  • Information and Cyber Security

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1007/s41060-025-00996-5

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.