Safety and efficiency are key challenges in autonomous vehicle platooning, particularly in scenarios where vehicles operate without inter-vehicle communication and must rely solely on local sensing. This work proposes a decentralized multi-agent platooning strategy based on Deep Reinforcement Learning, adopting a Predecessor-Based Sequential Training scheme in which follower vehicles are trained sequentially along the platoon. Under this approach, each agent learns a control policy using only local measurements of the preceding vehicle, enabling fully decentralized execution. The framework is evaluated in two representative scenarios: a highway environment with non-zero road grade and an urban setting characterized by stop-and-go leader behavior. Simulation results show that the proposed approach maintains safe inter-vehicle spacing and accurate speed tracking across both scenarios, demonstrating robustness to environmental disturbances and varying traffic conditions.
Sequential Multi-Agent Deep Reinforcement Learning for Platooning Across Diverse Road and Leader Behaviors / Berdini, E., Wrona, A., Menegatti, D., Delli Priscoli, F.. - (2026), pp. 183-188. (2026 34th Mediterranean Conference on Control and Automation (MED) Ancona; Italy ) [10.1109/med70602.2026.11598474].
Sequential Multi-Agent Deep Reinforcement Learning for Platooning Across Diverse Road and Leader Behaviors
Berdini, EmanueleSoftware
;Wrona, AndreaMethodology
;Menegatti, Danilo
Conceptualization
;Delli Priscoli, FrancescoFunding Acquisition
2026
Abstract
Safety and efficiency are key challenges in autonomous vehicle platooning, particularly in scenarios where vehicles operate without inter-vehicle communication and must rely solely on local sensing. This work proposes a decentralized multi-agent platooning strategy based on Deep Reinforcement Learning, adopting a Predecessor-Based Sequential Training scheme in which follower vehicles are trained sequentially along the platoon. Under this approach, each agent learns a control policy using only local measurements of the preceding vehicle, enabling fully decentralized execution. The framework is evaluated in two representative scenarios: a highway environment with non-zero road grade and an urban setting characterized by stop-and-go leader behavior. Simulation results show that the proposed approach maintains safe inter-vehicle spacing and accurate speed tracking across both scenarios, demonstrating robustness to environmental disturbances and varying traffic conditions.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


