MARATTO

article · ITM Web of Conferences

Optimizing Convolution Operations for YOLOv4-based Object Detection on GPU

2024Open accessMohammed V University

Abstract

Real-time object detection is crucial for autonomous vehicles, and YOLO (You Only Look Once) algorithms have demonstrated their effectiveness for this purpose. This study examines the performance of YOLOv4 [3] for real-time object detection on an embedded architecture. We focus on optimizing the computationally intensive convolution operations by employing the cuDNN library to achieve efficient inference. The evaluation assesses critical performance metrics, including object detection accuracy in terms of Mean Average Precision (mAP) and inference latency on the embedded architecture. We conduct a comparative analysis using the publicly available KITTI [7] database. The reported results establish a benchmark between the parallelized YOLOv4 model and the baseline implementation, assessing the advantages of cuDNN acceleration for real-time object detection on resource-constrained devices.

Research topics

  • Advanced Neural Network Applications
  • Robotics and Sensor-Based Localization
  • Advanced Image and Video Retrieval Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1051/itmconf/20246904008

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.