MARATTO

article · Statistics, Optimization & Information Computing (University of Cambridge)

Adaptive Operator Selection in ALNS Using Q-Learning: A Comparative Study of Softmax, UCB, epsilon-Greedy and QTS Strategies for CVRP

Abstract

This paper presents a comprehensive comparative study of operator selection strategies within a previously proposed Q-learning-based Adaptive Large Neighborhood Search (ALNS) framework for solving the Capacitated Vehicle Routing Problem (CVRP). Unlike classical ALNS approaches, where operators are selected using predefined heuristic rules, the proposed framework dynamically learns effective destroy–repair operator pairs during the search process. In addition, a new Q-Value Thompson Sampling (QTS) action selection strategy is introduced and compared with the conventional Roulette Wheel Selection baseline as well as three widely used reinforcement learning policies, namely ϵ-greedy, Softmax, and Upper Confidence Bound (UCB). The five strategies are evaluated using representative benchmark instances from CVRPLIB under identical experimental conditions over 30 independent runs. The comparative analysis considers solution quality, computational time, convergence behaviour, robustness through boxplot analysis, and statistical significance using Friedman and Wilcoxon signed-rank tests. The results show that the investigated strategies exhibit complementary strengths. These findings provide new insights into the impact of action selection policies on adaptive operator selection and demonstrate that QTS constitutes a robust and competitive alternative within the ALNS framework.

Research topics

  • Vehicle Routing Optimization Methods
  • Transportation Planning and Optimization
  • Traffic control and management

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.19139/soic-2310-5070-3956

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.