Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
Published in ACM Conference on Computer and Communications Security (CCS), 2026
Reinforcement learning enables adversaries to break safety alignment of LLMs more effectively than supervised fine-tuning. TokenBuncher is the first defense specifically targeting RL-based harmful fine-tuning: it suppresses model response entropy — the foundation RL exploits — through entropy-as-reward RL and a Token Noiser mechanism, preventing escalation of harmful capabilities while preserving benign utility.
Authors: Weitao Feng, Lixu Wang, Tianyi Wei, Jie Zhang, Chongyang Gao, Simon Sinong Zhan, Peizhuo Lv, Wei Dong
Citation
@inproceedings{feng2026token, title={Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning}, author={Feng, Weitao and Wang, Lixu and Wei, Tianyi and Zhang, Jie and Gao, Chongyang and Zhan, Sinong and Lv, Peizhuo and Dong, Wei}, booktitle={ACM Conference on Computer and Communications Security (CCS)}, year={2026}, url={https://arxiv.org/abs/2508.20697} }