article · FUDMA Journal of Sciences
Modern computer networks rely heavily on intrusion detection and prevention systems to defend against cyber threats. A systematic review of network security literature highlights critical issues in how these detection tools are tested. Widely used benchmark datasets, such as KDD Cup 99, remain prevalent despite containing duplicate records and obsolete attack profiles that fail to reflect current threats. Newer datasets offering realistic traffic patterns and specific Internet of Things scenarios, such as CICIDS2017 and Bot-IoT, still see comparatively modest uptake. Furthermore, most evaluations rely on artificial laboratory traffic rather than live operational environments, restricting the real-world validity of reported performance. This creates an ongoing mismatch between experimental benchmarks and operational conditions. To address this gap, guidance is offered to help select testing environments that improve the realism, dependability, and transferability of intrusion detection solutions.
Intrusion detection systems protect critical digital infrastructure from attacks, but their true effectiveness depends on realistic testing. When security tools are evaluated on outdated or synthetic network data, their reported success rates can be misleading. Demonstrating system performance in conditions that mirror modern operational traffic is essential to ensure that deployed cybersecurity tools reliably detect contemporary threats in practical environments.
This work offers benchmarking guidance for cybersecurity software developers, network security vendors, and testing laboratories seeking to validate intrusion detection products. Because it is a review of evaluation practices rather than a deployable tool, it sits at the methodological guidance stage. It can inform product testing protocols, enabling security teams to better simulate operational conditions and verify commercial readiness before deploying threat-detection solutions to live corporate or industrial networks.
AI-generated from the published abstract. Always read the original work before citing.
Most modern networks depend on Intrusion Detection Systems (IDS) and Intrusion Prevention Systems (IPS) as core defense system against an increasingly expanding number of cyber threats. This is the reason why research effort over the years has concentrated on improving these systems, while attention is now been directed at the network settings in which they are tested. In this paper, we systematically review the IDS/IPS literature and discuss the evaluation environments and benchmark datasets used to evaluate detection performance, e.g., KDD Cup 99, NSL-KDD, UNSW-NB15, CICIDS2017, and Bot-IoT. The analysis shows that the popular datasets like the KDD Cup 99 are still in widespread use despite their known limitations, such as duplicate records and attack profiles that do not reflect modern threats. Nevertheless, the adoption of other recent datasets such as CICIDS2017 and Bot-IoT that capture more realistic traffic, and include IoT-specific scenarios, is still modest compared to their older counterparts. The review also shows that experiments are mostly based on artificial and not operational traffic and are mostly limited to laboratory and not production settings, thus limiting the external validity of reported results. Taken together, these observations establish a persistent discrepancy between benchmark settings and the realities of operational networks. The paper ends with practical recommendations to assist researchers and practitioners to choose evaluation environments to enhance the realism, dependability and transferability of IDS/IPS solutions.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.33003/fjs-2026-1014-5147
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.