article · Sensors
Identifying violent behaviour across real-world closed-circuit television systems is difficult due to varying camera specifications and environments. To address this, a novel video violence detection model was developed using a three-dimensional convolutional neural network based on ResNet-3D architecture. The system processes high-dimensional video inputs by combining standard RGB visual data with optical flow to capture critical spatial and temporal movements. An integrated attention mechanism weights the most significant video frames during detected incidents, improving both detection accuracy and computational efficiency. When tested against standard benchmark datasets including UBI-Fight, Hockey, Crowd, and Movie-Fights, the architecture achieved area under the curve scores between 94.5 and 100.0, outperforming existing state-of-the-art methods across all evaluated scenarios.
Automating violence detection in video surveillance helps security operators identify dangerous incidents rapidly without manually monitoring numerous camera feeds. By successfully interpreting motion and visual context across diverse video conditions, this approach supports more reliable public safety monitoring and broadens capabilities in automated video analysis.
This technology is relevant to developers of public security infrastructure, surveillance software providers, and control room operators. While the model has been applied and tested successfully on several benchmark datasets and shows potential for integration into real-time closed-circuit television networks, further adaptation would be required to transition it from benchmark testing into operational commercial monitoring systems.
AI-generated from the published abstract. Always read the original work before citing.
Detecting violent behavior in videos to ensure public safety and security poses a significant challenge. Precisely identifying and categorizing instances of violence in real-life closed-circuit television, which vary across specifications and locations, requires comprehensive understanding and processing of the sequential information embedded in these videos. This study aims to introduce a model that adeptly grasps the spatiotemporal context of videos within diverse settings and specifications of violent scenarios. We propose a method to accurately capture spatiotemporal features linked to violent behaviors using optical flow and RGB data. The approach leverages a Conv3D-based ResNet-3D model as the foundational network, capable of handling high-dimensional video data. The efficiency and accuracy of violence detection are enhanced by integrating an attention mechanism, which assigns greater weight to the most crucial frames within the RGB and optical-flow sequences during instances of violence. Our model was evaluated on the UBI-Fight, Hockey, Crowd, and Movie-Fights datasets; the proposed method outperformed existing state-of-the-art techniques, achieving area under the curve scores of 95.4, 98.1, 94.5, and 100.0 on the respective datasets. Moreover, this research not only has the potential to be applied in real-time surveillance systems but also promises to contribute to a broader spectrum of research in video analysis and understanding.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.3390/s24020317
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.