MARATTO

book chapter · Advances in computational intelligence and robotics book series

Evaluating Pre-Trained CNNs and Vision Transformers in Image Classification

Abstract

This study compares pre-trained Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) for emergency vehicle classification using a curated dataset of 2312 images. Pre-processing involved standardizing image dimensions and parsing XML annotations. Google Colab's GPUs were used to train models with the Adam optimizer and sparse categorical cross-entropy loss. The results revealed that ViTs are competitive with CNNs in classification performance. Among CNNs, MobileNetV1 achieved the highest accuracy of 92.21% with a training time of 32.47 seconds. ViTs, including ViT_B_32 and ViT_B_16, reached accuracies of 95.83% and 93.75%, respectively, but required longer training times. These findings, demonstrate the strength of ViTs in handling long-range dependencies while remaining competitive with traditional CNNs. This research provides valuable insights into architecture selection for image classification, contributing to the advancement of deep learning in emergency vehicle systems.

Research topics

  • Advanced Neural Network Applications
  • Advanced Image and Video Retrieval Techniques
  • Brain Tumor Detection and Classification

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.4018/979-8-3373-1220-0.ch016

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.