MARATTO

article

Evaluating Machine Learning Models Best Fit for Crime Prediction in Windhoek

Abstract

Crime refers to any unlawful act which is punishable by law. Crime has a harmful effect on the community, it hinders the economic growth and it causes fears in the society. Crime rates have been on the increase around the world, which led to law enforcement agencies and stakeholders to implement crime prevention approaches such as predictive policing as most of the traditional crime-solving techniques were proven to be incompetent and ineffective. In Namibia, there is a lack of data-driven predictive method for combating crime. Hence, the paper evaluated machine learning (ML) models best fit for predicting crime types in Windhoek. The study employed a quantitative research approach to historical crime data from January 2018 to December 2022. After pre-processing, supervised classification models namely; Random Forest, K-Nearest Neighbors (KNN), Naive Bayes, and Decision Trees were trained and tested on the dataset. Initial findings revealed low accuracy rates across all models due to the imbalanced class distribution within the crime dataset mainly on the target variable which is Type of crime. To tackle the imbalances, various resampling techniques namely, Synthetic Minority Over-sampling Technique (SMOTE), Edited Nearest Neighbor (ENN), SMOTETomek and SMOTEEN were applied, leading to a significant improvement in model performance. The Random Forest model demonstrated superior performance, especially when trained on balanced datasets created through SMOTE and SMOTETomek, achieving an accuracy rate of 84%. The paper highlights the important role of addressing class imbalance through resampling techniques to improve the predictive accuracy of machine learning models in crime prediction. The findings suggest that the Random Forest model, coupled with appropriate resampling methods, provides a robust solution for crime type prediction in Windhoek.

Research topics

  • Crime Patterns and Interventions
  • Big Data Technologies and Applications
  • Data Mining and Machine Learning Applications

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icabcd62167.2024.10645286

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.