article · Scientific Reports
Internet of Things-based wireless sensor networks (IoT-WSNs) face persistent challenges related to energy consumption, latency, and network congestion under dynamic and heterogeneous topologies. Conventional reinforcement learning approaches rely on static reward formulations, which limit adaptability and hinder effective multi-objective optimization. This study proposes a dynamic reward structuring framework within deep reinforcement learning to enable adaptive and balanced routing in IoT-WSNs. The proposed approach employs real-time reward recalibration to jointly optimize energy efficiency, delay, and throughput under varying network conditions. A hybrid deep reinforcement learning architecture is developed by integrating value-based, policy-based, and actor-critic methods, along with multi-agent coordination and attention mechanisms to prioritize critical nodes and links. Furthermore, a hierarchical learning structure decomposes global and local routing objectives, improving scalability and decision efficiency in complex network environments. Experimental results demonstrate that the proposed framework achieves significant performance gains, including approximately 30% improvement in energy efficiency, 25% reduction in latency, and 35% increase in network throughput compared with baseline methods. These findings highlight the effectiveness of dynamic reward adaptation for scalable and robust multi-objective optimization in IoT-WSN routing.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1038/s41598-026-61080-x
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.