hierarchical reinforcement learning for multipath flow
Topology-aware hierarchical reinforcement learning for multipath commodity flow optimization.
question
How can a network construct multiple routing paths and allocate traffic among them in response to real-time load, while satisfying capacity and routing constraints?
motivation
Many approaches to the multi-path commodity flow problem rely on pre-generated candidate paths. This separates path construction from traffic allocation and makes it difficult to respond to changing network states or account for the strong coupling between the two decisions.
approach
We propose HRL-TAE, an end-to-end hierarchical reinforcement learning framework with topology-aware exploration.
- A State Transition Guiding List turns topology exploration into a structured Markov decision process.
- A path construction layer builds multiple subpaths for each flow.
- A ratio allocation layer distributes traffic across the constructed paths.
- Adaptive constraint-driven masks remove infeasible actions and ensure that generated solutions satisfy the problem constraints.
- A tailored training strategy improves gradient estimation and training efficiency.
findings
Experiments in simulation and real-world network settings show that HRL-TAE can dynamically produce high-quality feasible solutions and outperform the evaluated baselines.
publication
Hierarchical Reinforcement Learning with Topology-Aware Exploration Framework for Multi-path Commodity Flow Problem
Jingchen Jiang, Xuan Zhou, Jiayuan Li, Geng Han, Xiang Shi, and Fang Deng
Proceedings of the AAAI Conference on Artificial Intelligence, 40(43), 36280–36288, 2026.
Paper page · PDF · DOI