Posts by Collection

portfolio

publications

RElectrode: A Reconfigurable Electrode For Multi-Purpose Sensing Based on Microfluidics

Published in ACM Conference on Human Factors in Computing Systems (CHI), 2021

RElectrode is a reconfigurable electrode using a microfluidic technique that can change the geometry and material properties of the electrode to satisfy the needs for sensing a variety of different types of user input through touch/touchless gestures, pressure, temperature, and distinguish between different types of objects or liquids.

Recommended citation: @inproceedings{sun2021relectrode, title={RElectrode: A Reconfigurable Electrode For Multi-Purpose Sensing Based on Microfluidics}, author={Sun, Wei and Chen, Yanjun and Zhan, Simon and Han, Teng and Tian, Feng and Wang, Hongan and Yang, Xing-Dong}, booktitle={Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems}, year={2021}, doi={10.1145/3411764.3445652}, url={https://doi.org/10.1145/3411764.3445652} }
Download Paper

MicroFluID - A Reconfigurable RFID Platform for Robust Interaction Sensing Based on Microfluidics

Published in ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp), 2022

MicroFluID is a novel RFID artifact based on a multiple-chip structure and microfluidic switches, which informs the input state by directly reading variable ID information instead of retrieving primitive signals.

Recommended citation: @article{sun2022microfluid, title={MicroFluID - A Reconfigurable RFID Platform for Robust Interaction Sensing Based on Microfluidics}, author={Sun, Wei and Chen, Yuwen and Chen, Yanjun and Zhang, Xiaopeng and Zhan, Simon and Li, Yixin and Wu, Jiecheng and Han, Teng and Mi, Haipeng and Wang, Jingxian and Tian, Feng and Yang, Xing-Dong}, journal={Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies}, volume={6}, number={3}, year={2022}, doi={10.1145/3550296}, url={https://dl.acm.org/doi/abs/10.1145/3550296} }
Download Paper

Joint Differentiable Optimization and Verification for Certified Reinforcement Learning

Published in ACM/IEEE International Conference on Cyber-Physical Systems (ICCPS), 2023

A framework that jointly conducts reinforcement learning and formal verification by formulating and solving a novel bilevel optimization problem, which is end-to-end differentiable by the gradients from the value function and certificates formulated by linear programs and semi-definite programs.

Recommended citation: @inproceedings{wang2023joint, title={Joint differentiable optimization and verification for certified reinforcement learning}, author={Wang, Yixuan and Zhan, Simon and Wang, Zhilu and Huang, Chao and Wang, Zhaoran and Yang, Zhuoran and Zhu, Qi}, booktitle={Proceedings of the ACM/IEEE 14th International Conference on Cyber-Physical Systems (with CPS-IoT Week 2023)}, pages={132--141}, year={2023} }
Download Paper

Enforcing Hard Constraints with Soft Barriers: Safe Reinforcement Learning in Unknown Stochastic Environments

Published in International Conference on Machine Learning (ICML), 2023

A safe RL approach that can jointly learn the environment and optimize the control policy, while effectively avoiding unsafe regions with safety probability optimization.

Recommended citation: @inproceedings{wang2023enforcing, title={Enforcing hard constraints with soft barriers: Safe reinforcement learning in unknown stochastic environments}, author={Wang, Yixuan and Zhan, Simon Sinong and Jiao, Ruochen and Wang, Zhilu and Jin, Wanxin and Yang, Zhuoran and Wang, Zhaoran and Huang, Chao and Zhu, Qi}, booktitle={International Conference on Machine Learning}, pages={36593--36604}, year={2023}, organization={PMLR} }
Download Paper

Empowering Autonomous Driving with Large Language Models: A Safety Perspective

Published in LLMAgent Workshop at ICLR 2024, 2024

This paper explores the integration of Large Language Models (LLMs) into autonomous driving systems, leveraging their robust common-sense knowledge and reasoning abilities to enhance driving performance and safety in long-tail unforeseen scenarios.

Recommended citation: @inproceedings{wang2024empowering, title={Empowering Autonomous Driving with Large Language Models: A Safety Perspective}, author={Wang, Yixuan and Jiao, Ruochen and Zhan, Simon and Lang, Chengtian and Huang, Chao and Wang, Zhaoran and Yang, Zhuoran and Zhu, Qi}, booktitle={LLMAgent Workshop at ICLR 2024}, year={2024}, url={https://arxiv.org/abs/2312.00812} }
Download Paper

State-wise Safe Reinforcement Learning With Pixel Observations

Published in Learning for Dynamics and Control Conference (L4DC), 2024

In this paper, we propose a novel pixel-observation safe RL algorithm that efficiently encodes state-wise safety constraints with unknown hazard regions through the introduction of a latent barrier function learning mechanism.

Recommended citation: @inproceedings{zhan2024statewise, title={State-wise Safe Reinforcement Learning With Pixel Observations}, author={Zhan, Sinong Simon and Wang, Yixuan and Wu, Qingyuan and Jiao, Ruochen and Huang, Chao and Zhu, Qi}, booktitle={Learning for Dynamics and Control Conference (L4DC)}, year={2024}, url={https://arxiv.org/abs/2311.02227} }
Download Paper

Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays

Published in International Conference on Machine Learning (ICML), 2024

Auxiliary-Delayed Reinforcement Learning (AD-RL) leverages an auxiliary short-delayed task to accelerate the learning on a strongly delayed task without compromising the performance in stochastic environments.

Recommended citation: @inproceedings{wu2024boosting, title={Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays}, author={Wu, Qingyuan and Zhan, Sinong Simon and Wang, Yixuan and Wang, Yuhui and Lin, Chung-Wei and Lv, Chen and Zhu, Qi and Schmidhuber, Jürgen and Huang, Chao}, booktitle={International Conference on Machine Learning (ICML)}, year={2024}, url={https://arxiv.org/abs/2402.03141} }
Download Paper

Kinematics-aware Trajectory Generation and Prediction with Latent SDE

Published in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024

This paper presents a novel approach to trajectory generation and prediction that incorporates kinematic constraints through latent stochastic differential equations, enabling more realistic and physically-consistent motion planning.

Recommended citation: @inproceedings{zhan2024kinematics, title={Kinematics-aware Trajectory Generation and Prediction with Latent SDE}, author={Zhan, Sinong Simon and Wu, Qingyuan and Wang, Yixuan and Huang, Chao and Zhu, Qi}, booktitle={IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)}, year={2024}, url={https://arxiv.org/abs/2309.09317} }
Download Paper

Switching Controller Synthesis for Hybrid Systems Against STL Formulas

Published in International Symposium on Formal Methods (FM), 2024

A sound and relatively complete approach for synthesizing switching controllers for hybrid systems against Signal Temporal Logic specifications, by iteratively computing state sets that satisfy reach-avoid objectives with timing constraints.

Recommended citation: @inproceedings{su2024switching, title={Switching Controller Synthesis for Hybrid Systems Against STL Formulas}, author={Su, Han and Feng, Shenghua and Zhan, Sinong and Zhan, Naijun}, booktitle={International Symposium on Formal Methods (FM)}, year={2024}, url={https://arxiv.org/abs/2406.16588} }
Download Paper

Variational Delayed Policy Optimization

Published in Conference on Neural Information Processing Systems (NeurIPS), 2024

Variational Delayed Policy Optimization (VDPO) reformulates delayed RL as a variational inference problem, which is further modelled as a two-step iterative optimization problem, where the first step is TD learning in the delay-free environment with a small state space, and the second step is behaviour cloning which can be addressed much more efficiently than TD learning.

Recommended citation: @article{wu2024variational, title={Variational delayed policy optimization}, author={Wu, Qingyuan and Zhan, Simon S and Wang, Yixuan and Wang, Yuhui and Lin, Chung-Wei and Lv, Chen and Zhu, Qi and Huang, Chao}, journal={Advances in neural information processing systems}, volume={37}, pages={54330--54356}, year={2024} }
Download Paper

Case Study: Runtime Safety Verification of Neural Network Controlled System

Published in International Conference on Runtime Verification (RV), 2024

A case study on using POLAR-Express, a state-of-the-art NNCS reachability analysis tool, for runtime safety verification in a Turtlebot navigation system with LiDAR observations.

Recommended citation: @inproceedings{yang2024case, title={Case Study: Runtime Safety Verification of Neural Network Controlled System}, author={Yang, Frank and Zhan, Simon Sinong and Wang, Yixuan and Huang, Chao and Zhu, Qi}, booktitle={International Conference on Runtime Verification (RV)}, year={2024}, url={https://arxiv.org/abs/2408.08592} }
Download Paper

Inverse Delayed Reinforcement Learning

Published in arXiv preprint (under review), 2024

An IRL framework that extracts rewarding features from expert trajectories affected by delayed disturbances, using an efficient off-policy adversarial training scheme to recover optimal policies from augmented delayed observations.

Recommended citation: @article{zhan2024inverse, title={Inverse Delayed Reinforcement Learning}, author={Zhan, Simon Sinong and Wu, Qingyuan and Ruan, Zhian and Yang, Frank and Wang, Philip and Wang, Yixuan and Jiao, Ruochen and Huang, Chao and Zhu, Qi}, journal={arXiv preprint arXiv:2412.02931}, year={2024} }
Download Paper

Directly Forecasting Belief for Reinforcement Learning with Delays

Published in International Conference on Machine Learning (ICML), 2025

This paper presents a novel approach to directly forecast beliefs in reinforcement learning with observation delays, improving upon traditional methods by incorporating predictive capabilities into the learning process.

Recommended citation: @inproceedings{zhan2025directly, title={Directly Forecasting Belief for Reinforcement Learning with Delays}, author={Zhan, Sinong Simon and Wu, Qingyuan and Wang, Yixuan and Huang, Chao and Zhu, Qi}, booktitle={International Conference on Machine Learning (ICML)}, year={2025}, url={https://arxiv.org/abs/2505.00546} }
Download Paper

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

Published in arXiv preprint (under review), 2025

SENTINEL grounds practical safety requirements of foundation-model-based embodied agents in formal temporal logic semantics and evaluates them at the semantic, plan, and trajectory levels within a unified formal framework.

Recommended citation: @article{zhan2025sentinel, title={SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents}, author={Zhan, Simon Sinong and Liu, Yao and Wang, Philip and Wang, Zinan and Wang, Qineng and Peng, Yiyan and Ruan, Zhian and Shi, Xiangyu and Cao, Xinyu and Yang, Frank and Wang, Kangrui and Shao, Huajie and Li, Manling and Zhu, Qi}, journal={arXiv preprint arXiv:2510.12985}, year={2025} }
Download Paper

See, Think, Act: Online Shopper Behavior Simulation with VLM Agents

Published in Scaling Environments for Agents (SEA) Workshop at NeurIPS 2025, 2025

The first systematic integration of textual context and visual webpage perception for online shopper behavior simulation, leveraging vision-language models to align agent decision-making with realistic human shopping patterns.

Recommended citation: @inproceedings{zhang2025see, title={See, Think, Act: Online Shopper Behavior Simulation with VLM Agents}, author={Zhang, Yimeng and Gesi, Jiri and Xue, Ran and Wang, Tian and Wang, Ziyi and Lu, Yuxuan and Zhan, Sinong and Zeng, Huimin and Cui, Qingjun and Guo, Yufan and Huang, Jing and Shah, Mubarak and Wang, Dakuo}, booktitle={Scaling Environments for Agents (SEA) Workshop at NeurIPS 2025}, year={2025}, url={https://arxiv.org/abs/2510.19245} }
Download Paper

Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning

Published in ICLR 2026/Scaling Environments for Agents (SEA) Workshop at NeurIPS 2025, 2026

This paper introduces Shop-R1, a novel reinforcement learning framework aimed at enhancing the reasoning ability of LLMs for simulation of real human behavior in online shopping environments through a two-stage approach with distinct reward signals.

Recommended citation: @inproceedings{zhang2025shop, title={Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning}, author={Zhang, Yimeng and Wang, Tian and Gesi, Jiri and Wang, Ziyi and Lu, Yuxuan and Lin, Jiacheng and Zhan, Sinong and Gao, Vianne and Jiao, Ruochen and Liu, Junze and Qian, Kun and Tang, Yuxin and Xue, Ran and Zhang, Houyu and Cui, Qingjun and Guo, Yufan and Wang, Dakuo}, booktitle={Scaling Environments for Agents (SEA) Workshop at NeurIPS 2025}, year={2025}, url={https://arxiv.org/abs/2507.17842} }
Download Paper

Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization

Published in ICLR 2026, 2026

DT-CORL learns delay-robust policies from static, delay-free offline data by jointly optimizing a transformer-based belief model and a constrained policy objective.

Recommended citation: @article{zhan2025adapting, title={Adapting Offline Reinforcement Learning with Online Delays}, author={Zhan, Simon Sinong and Wu, Qingyuan and Yang, Frank and Shi, Xiangyu and Huang, Chao and Zhu, Qi}, journal={arXiv preprint arXiv:2506.00131}, year={2025} }
Download Paper

Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping

Published in L4DC 2026, 2026

This paper proposes a transition-aware reward shaping framework for adversarial inverse reinforcement learning in stochastic environments, integrating transition model estimation to learn stochastic-invariant rewards and improve sample efficiency and performance.

Recommended citation: @article{zhan2024enhancing, title={Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping}, author={Zhan, Simon Sinong and Wang, Philip and Wu, Qingyuan and Jiao, Ruochen and Wang, Yixuan and Huang, Chao and Zhu, Qi}, journal={arXiv preprint arXiv:2410.03847}, year={2024} }
Download Paper

A Unified Framework for Rethinking Policy Divergence Measures in GRPO

Published in arXiv preprint (under review), 2026

A unified clipping framework that characterizes GRPO-style RLVR methods via a general notion of policy divergence — spanning likelihood ratios and KL divergences — and identifies the variance-reduced KL3 estimator as a key constraint.

Recommended citation: @article{wu2026unified, title={A Unified Framework for Rethinking Policy Divergence Measures in GRPO}, author={Wu, Qingyuan and Wang, Yuhui and Zhan, Simon Sinong and Dai, Yanning and Deng, Shilong and Habchi, Sarra and Zhu, Qi and Gallé, Matthias and Huang, Chao}, journal={arXiv preprint arXiv:2602.05494}, year={2026} }
Download Paper

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

Published in ACM Conference on Computer and Communications Security (CCS), 2026

TokenBuncher is the first effective defense against RL-based harmful fine-tuning of LLMs, suppressing response entropy via entropy-as-reward RL and a Token Noiser mechanism to prevent escalation of harmful capabilities.

Recommended citation: @inproceedings{feng2026token, title={Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning}, author={Feng, Weitao and Wang, Lixu and Wei, Tianyi and Zhang, Jie and Gao, Chongyang and Zhan, Sinong and Lv, Peizhuo and Dong, Wei}, booktitle={ACM Conference on Computer and Communications Security (CCS)}, year={2026}, url={https://arxiv.org/abs/2508.20697} }
Download Paper

STEP-LLM: Generating CAD STEP Models from Natural Language with Large Language Models

Published in Design, Automation and Test in Europe Conference (DATE), 2026

STEP-LLM is the first unified framework for direct STEP (ISO 10303) CAD file generation from natural language, combining a 40K STEP-caption dataset, graph-aware reserialization, retrieval-augmented fine-tuning, and Chamfer-distance-based RL refinement.

Recommended citation: @inproceedings{shi2026stepllm, title={STEP-LLM: Generating CAD STEP Models from Natural Language with Large Language Models}, author={Shi, Xiangyu and Ding, Junyang and Zhao, Xu and Zhan, Sinong and Mohapatra, Payal and Quispe, Daniel and Welbeck, Kojo and Cao, Jian and Chen, Wei and Guo, Ping and Zhu, Qi}, booktitle={Design, Automation and Test in Europe Conference (DATE)}, year={2026}, url={https://arxiv.org/abs/2601.12641} }
Download Paper

Shedding Light on VLN Robustness: A Black-box Framework for Indoor Lighting-based Adversarial Attack

Published in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

An indoor lighting-based adversarial attack revealing robustness gaps of vision-and-language navigation agents under realistic illumination perturbations.

Recommended citation: @inproceedings{li2026shedding, title={Shedding Light on VLN Robustness: A Black-box Framework for Indoor Lighting-based Adversarial Attack}, author={Li, Chenyang and Tang, Wenbing and Huang, Yihao and Zhan, Sinong Simon and Hu, Ming and Jia, Xiaojun and Liu, Yang}, booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year={2026}, url={https://arxiv.org/abs/2511.13132} }
Download Paper

talks

teaching