A Unified Framework for Rethinking Policy Divergence Measures in GRPO
Published in arXiv preprint (under review), 2026
Reinforcement Learning with Verified Reward (RLVR) is a central paradigm for advancing LLM reasoning. This work introduces a unified clipping framework that characterizes existing GRPO-style methods through a general notion of policy divergence, encompassing both likelihood ratios and KL divergences and extending to alternative measures, and identifies the variance-reduced KL3 Monte Carlo estimator as a key policy divergence constraint.
Authors: Qingyuan Wu, Yuhui Wang, Simon Sinong Zhan, Yanning Dai, Shilong Deng, Sarra Habchi, Qi Zhu, Matthias Gallé, Chao Huang
Citation
@article{wu2026unified, title={A Unified Framework for Rethinking Policy Divergence Measures in GRPO}, author={Wu, Qingyuan and Wang, Yuhui and Zhan, Simon Sinong and Dai, Yanning and Deng, Shilong and Habchi, Sarra and Zhu, Qi and Gallé, Matthias and Huang, Chao}, journal={arXiv preprint arXiv:2602.05494}, year={2026} }