MARATTO

article

Generalizable Conditional Imitation Learning for Urban Driving: A Multimodal ConvLSTM Approach in Adverse Conditions

Abstract

End-to-end learning for autonomous driving offers the potential to directly map sensory observations to control actions; however, such approaches often suffer from high sample complexity and poor generalization in urban environments. Conditional Imitation Learning (CIL) mitigates these limitations by conditioning the driving policy on high-level navigation commands, enabling more structured and goal-directed decision making. In this work, we present a robust and generalizable CIL framework based on a multimodal ConvLSTM architecture that fuses spatial–temporal features from RGB images, semantic segmentation, vehicle speed, traffic light state, and route instructions. The proposed model is evaluated in the CARLA simulator across three trajectories under dynamic traffic conditions involving vehicles and pedestrians. Experimental results demonstrate that the CIL-LSTM model achieves a 94% success rate, reduces collisions with road users to 1% (compared to 2% with DQN), and decreases route completion time by an average of 5.3 s, while maintaining comparable trajectory distances. These findings highlight the effectiveness of temporal modeling and multimodal perception in improving the robustness, safety, and scalability of autonomous driving systems.

Research topics

  • Multimodal Machine Learning Applications
  • Reinforcement Learning in Robotics
  • Domain Adaptation and Few-Shot Learning

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icateee68170.2025.11406391

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.