Authors:

  • Milind Tambe
This paper introduces a novel reinforcement learning framework designed for environments with combinatorial action spaces. The authors propose a latent spherical flow policy that leverages diffusion-based generative modeling techniques to efficiently represent and optimize complex action combinations. The approach improves exploration and policy learning in high-dimensional combinatorial settings, demonstrating strong empirical performance across benchmark reinforcement learning tasks. The work contributes new methods for scalable decision-making in domains where agents must select structured sets of actions simultaneously.

Citations

Kong L, Satish A, Jiang H, Kangaslahti A, Ma A, Chen W, Song M, Xu L, Tambe M. Latent spherical flow policy for reinforcement learning with combinatorial actions. In: Proceedings of the International Conference on Machine Learning (ICML). 2026. Spotlight paper.