ManiGuard: A Benchmark and Data Suite for Specification-Grounded Safety Evaluation and Improvement of Robotic Manipulation
Published in arXiv preprint, 2026
Foundation-model policies for robotic manipulation are advancing rapidly on task success, but rigorous evaluation of whether they succeed safely is still lacking. ManiGuard is a specification-grounded framework for evaluating and improving the safety of foundation-model manipulation. ManiGuard-Bench organizes six contact-rich household task families into 200 base tasks along a skill × constraint taxonomy, with safety specified independently of task success, and evaluates each under one in-distribution and four out-of-distribution perturbations (1,000 scenarios). A paired trajectory-generation pipeline provides 8,000 safety-annotated demonstrations for improving policy safety.
Authors: Yiyan Peng, Philip Wang, Simon Sinong Zhan, Yiqi Lyu, Zhenyang Ni, Jixin Yan, Fiorelli Wong, Ruochen Jiao, Hang Yin, Xinyu Cao, Huajie Shao, Manling Li, Ruohan Zhang, Qi Zhu (equal contribution)
Citation
@misc{peng2026maniguard, title={ManiGuard: A Benchmark and Data Suite for Specification-Grounded Safety Evaluation and Improvement of Robotic Manipulation}, author={Peng, Yiyan and Wang, Philip and Zhan, Simon Sinong and Lyu, Yiqi and Ni, Zhenyang and Yan, Jixin and Wong, Fiorelli and Jiao, Ruochen and Yin, Hang and Cao, Xinyu and Shao, Huajie and Li, Manling and Zhang, Ruohan and Zhu, Qi}, year={2026}, eprint={2608.17386}, archivePrefix={arXiv}, primaryClass={cs.RO}, url={https://arxiv.org/abs/2608.17386} }