article · Statistics, Optimization & Information Computing (University of Cambridge)
This paper presents a comprehensive comparative study of operator selection strategies within a previously proposed Q-learning-based Adaptive Large Neighborhood Search (ALNS) framework for solving the Capacitated Vehicle Routing Problem (CVRP). Unlike classical ALNS approaches, where operators are selected using predefined heuristic rules, the proposed framework dynamically learns effective destroy–repair operator pairs during the search process. In addition, a new Q-Value Thompson Sampling (QTS) action selection strategy is introduced and compared with the conventional Roulette Wheel Selection baseline as well as three widely used reinforcement learning policies, namely ϵ-greedy, Softmax, and Upper Confidence Bound (UCB). The five strategies are evaluated using representative benchmark instances from CVRPLIB under identical experimental conditions over 30 independent runs. The comparative analysis considers solution quality, computational time, convergence behaviour, robustness through boxplot analysis, and statistical significance using Friedman and Wilcoxon signed-rank tests. The results show that the investigated strategies exhibit complementary strengths. These findings provide new insights into the impact of action selection policies on adaptive operator selection and demonstrate that QTS constitutes a robust and competitive alternative within the ALNS framework.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.19139/soic-2310-5070-3956
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.