MARATTO

article · Scientific Reports

Multi-objective inventory optimization using reinforcement learning: a comparative study on profitability and carbon emissions

2026Open accessAin Shams University

Abstract

Inventory management is a core part of supply chains, and over the years it has been increasingly challenged by the need to balance economic performance with environmental considerations. While prior reinforcement learning (RL) studies have incorporated carbon emissions indirectly through cost penalties or regulatory constraints, this work addresses an existing gap by treating emissions as an independent optimization objective. This study examines RL as an adaptive decision‑making approach for inventory optimization with two objectives: maximizing profit and minimizing carbon emissions. The problem is formulated as a Markov Decision Process, and four RL algorithms Proximal Policy Optimization (PPO), Phasic Policy Gradient (PPG), Advantage Actor‑Critic (A2C), and Double Deep Q‑Network (DDQN) are evaluated under identical experimental conditions. Carbon emissions are explicitly modeled in the reward function rather than embedded within operating costs. The results show that PPG achieves the highest profitability with only a modest increase in emissions, while DDQN converges faster but yields lower profit overall. Sensitivity analysis indicates that reward weighting strongly influences policy behavior, with PPO providing the most stable trade‑off between profitability and emissions.

Research topics

  • Supply Chain and Inventory Management
  • Sustainable Supply Chain Management
  • Vehicle Routing Optimization Methods

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1038/s41598-026-44293-y

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.