See, Think, Act: Online Shopper Behavior Simulation with VLM Agents

Published in Scaling Environments for Agents (SEA) Workshop at NeurIPS 2025, 2025

This work investigates integrating visual information — webpage screenshots — into human shopping behavior simulation via vision-language models, using the publicly available OPeRA dataset. Each input instance combines the current webpage screenshot, full action history, and pruned HTML observations within the same session, aligning agent decision-making with realistic human online shopping patterns.

Authors: Yimeng Zhang, Jiri Gesi, Ran Xue, Tian Wang, Ziyi Wang, Yuxuan Lu, Sinong Zhan, Huimin Zeng, Qingjun Cui, Yufan Guo, Jing Huang, Mubarak Shah, Dakuo Wang

Citation

@inproceedings{zhang2025see, title={See, Think, Act: Online Shopper Behavior Simulation with VLM Agents}, author={Zhang, Yimeng and Gesi, Jiri and Xue, Ran and Wang, Tian and Wang, Ziyi and Lu, Yuxuan and Zhan, Sinong and Zeng, Huimin and Cui, Qingjun and Guo, Yufan and Huang, Jing and Shah, Mubarak and Wang, Dakuo}, booktitle={Scaling Environments for Agents (SEA) Workshop at NeurIPS 2025}, year={2025}, url={https://arxiv.org/abs/2510.19245} }