International Conference on Machine Learning (ICML)
Date of Publication:
2026
This paper studies inference-time alignment in large-scale AI systems through the lens of Stackelberg game theory. The authors propose a reward-shaping framework that models alignment as a strategic interaction between a leader and follower, enabling more robust and controllable decision-making during inference. The approach introduces theoretically grounded reward modifications that improve alignment performance while preserving policy effectiveness. Experimental evaluations demonstrate that the method enhances alignment robustness across reinforcement learning and generative AI settings.
Citations
Wang H, Lin T, Kong L, Li C, Jiang H, Tambe M. Reward shaping for (inference-time) alignment: a Stackelberg game perspective. In: Proceedings of the International Conference on Machine Learning (ICML). 2026.