I am a Ph.D. student at Shanghai Jiao Tong University (SJTU), advised by Prof. Dahua Lin. I conduct my research at Shanghai AI Laboratory, where I am jointly advised by Yuhang Zang and Jiaqi Wang.

My current research focuses on recursive self-improvement (RSI), agent harness engineering, and agentic post-training. I am interested in building agents that can continually improve through interaction, stronger training environments, and scalable feedback. I welcome discussions and collaborations in these areas.

Recursive Self-Improvement Agent Harness Engineering Agentic Post-Training Multimodal Agents

📢 News

SEAgent is accepted by ICML 2026.

We release WildClawBench, a benchmark for real-world, long-horizon agent evaluation.

We release Visual-ERM and VisualCritic-RewardBench.

ARM-Thinker is accepted by CVPR 2026.

RAR is accepted by IEEE Transactions on Image Processing.

We release SPARK, a policy and reward co-evolving framework.

We release CODA, a dual-brain computer-use agent.

Earlier news

Visual-RFT is accepted by ICCV 2025.

InternLM-XComposer2.5-Reward is accepted by Findings of ACL 2025.

MIA-DPO is accepted by ICLR 2025.

MMDU and MMLongBench-Doc (Spotlight) are accepted by NeurIPS 2024.

🔬 Publications

🤖 Agents

2026

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation PaperWildClawBench GitHub stars

Shuangrui Ding*, Xuanlang Dai*, Long Xing*, Shengyuan Ding, Ziyu Liu, Jingyi Yang, Penghui Yang, Zhixiong Zhang, Xilin Wei, Xinyu Fang, et al.

An in-the-wild benchmark for evaluating long-horizon agents in realistic OpenClaw environments.

ICML 2026

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience PaperSEAgent GitHub stars

Zeyi Sun, Ziyu Liu, Yuhang Zang, Yuhang Cao, Xiaoyi Dong, Tong Wu, Dahua Lin, Jiaqi Wang

A computer-use agent that autonomously learns novel software through curriculum generation, experience, and reinforcement learning.

2025

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning PaperCODA GitHub stars

Zeyi Sun, Yuhang Cao, Jianze Liang, Qiushi Sun, Ziyu Liu, Zhixiong Zhang, Yuhang Zang, Xiaoyi Dong, Kai Chen, Dahua Lin, Jiaqi Wang

A trainable planner-executor framework for robust execution and generalization in scientific computing GUIs.

2025

Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning First Author PaperVisual-RFT GitHub stars

Ziyu Liu, Yuhang Zang, Yushan Zou, Zijian Liang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang

Reinforcement fine-tuning that equips open-source LVLMs with web search and visual coding capabilities.

🎯 RL & Reward Models

2026

Visual-ERM: Reward Modeling for Visual Equivalence First Author PaperVisual-ERM GitHub stars

Ziyu Liu, Shengyuan Ding, Xinyu Fang, Xuanlang Dai, Penghui Yang, Jianze Liang, Jiaqi Wang, Kai Chen, Dahua Lin, Yuhang Zang

A multimodal generative reward model that evaluates vision-to-code outputs in rendered visual space with fine-grained, interpretable feedback.

CVPR 2026

ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning PaperARM-Thinker GitHub stars

Shengyuan Ding, Xinyu Fang, Ziyu Liu, Yuhang Zang, Yuhang Cao, Xiangyu Zhao, Haodong Duan, Xiaoyi Dong, Jianze Liang, Bin Wang, Conghui He, Dahua Lin, Jiaqi Wang

A multimodal reward model that reasons over visual inputs and actively uses tools to provide stronger feedback.

2025

SPARK: Synergistic Policy And Reward Co-Evolving Framework First Author PaperSPARK GitHub stars

Ziyu Liu, Yuhang Zang, Shengyuan Ding, Yuhang Cao, Xiaoyi Dong, Haodong Duan, Dahua Lin, Jiaqi Wang

An on-policy framework in which policy and generative reward modeling improve together using recycled rollout supervision.

ICCV 2025

Visual-RFT: Visual Reinforcement Fine-Tuning First Author PaperVisual-RFT GitHub stars

Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang

The first comprehensive adaptation of R1-style reinforcement fine-tuning to multimodal perception tasks with verifiable rewards.

ACL 2025

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model PaperInternLM-XComposer GitHub stars

Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Ziyu Liu, Shengyuan Ding, Shenxi Wu, Yubo Ma, Haodong Duan, Wenwei Zhang, Kai Chen, Dahua Lin, Jiaqi Wang

A general multimodal reward model for RL supervision, test-time response selection, and training-data filtering.

ICLR 2025

MIA-DPO: Multi-Image Augmented Direct Preference Optimization for Large Vision-Language Models First Author PaperMIA-DPO GitHub starsProject

Ziyu Liu, Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Haodong Duan, Conghui He, Yuanjun Xiong, Dahua Lin, Jiaqi Wang

A visual preference alignment method that extends DPO to diverse multi-image inputs without costly new annotations.

📚 Other Publications

TIP 2025

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition First Author PaperRAR GitHub starsProject

Ziyu Liu, Zeyi Sun, Yuhang Zang, Wei Li, Pan Zhang, Xiaoyi Dong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang

Retrieval and ranking augment multimodal language models for fine-grained, few-shot, and zero-shot visual recognition.

NeurIPS 2024

MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs First Author PaperMMDU GitHub starsProject

Ziyu Liu, Tao Chu, Yuhang Zang, Xilin Wei, Xiaoyi Dong, Pan Zhang, Zijian Liang, Yuanjun Xiong, Yu Qiao, Dahua Lin, Jiaqi Wang

A benchmark and 45K instruction-tuning dataset for long-context, multi-turn, multi-image conversations.

NeurIPS 2024 Spotlight

MMLongBench-Doc: Benchmarking Long-Context Document Understanding with Visualizations PaperMMLongBench-Doc GitHub stars

Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, Pan Zhang, Liangming Pan, Yu-Gang Jiang, Jiaqi Wang, Yixin Cao, Aixin Sun

A benchmark for understanding long documents containing text, tables, charts, images, and page layouts.

📝 Technical Reports

2024-2026

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models ReportVLMEvalKit GitHub stars

An open-source toolkit for reproducible evaluation of large multimodality models.

2026

Intern-S1-Pro: A Trillion-Scale Multimodal Foundation Model for Scientific Reasoning ReportModelIntern-S GitHub stars

A trillion-scale multimodal foundation model for scientific reasoning.

2026

Intern-S2-Preview ModelIntern-S GitHub stars

A scientific multimodal foundation model with stronger reasoning and long-horizon agent capabilities.

🛠️ Open Source

Selected developer tools and Agent Skills. Research code is linked with each publication above.

Developer Tool · 2026

ClaudeScope Project Lead ProjectClaudeScope GitHub stars

Parse and visualize Claude Code session trajectories, token usage, tool calls, timing, and failures.

Agent Skill · 2026

gene.skill Project Lead Projectgene.skill GitHub stars

Extract, recombine, and mutate capability genes from existing Agent Skills to create new composite skills.

Agent Skill · 2026

AnythingAtlas Project Lead AnythingAtlas GitHub stars

Turn scattered resources into a curated map and personalized learning path for any unfamiliar topic.

RAG System · 2024

SODA Project Lead SODA GitHub stars

Search, organize, and discover information across the web and private local knowledge bases.

🏆 Honors and Awards

  • 2025.09, Zhang Liangqi Scholarship, Shanghai Jiao Tong University.
  • 2024.06, Excellent Bachelor’s Thesis and Outstanding Undergraduate Graduate.
  • 2023.07, Meritorious Award, Mathematical Contest in Modeling and Interdisciplinary Contest in Modeling (MCM/ICM), COMAP. View certificate (PDF)
  • 2022.05, National Second Prize, China Undergraduate Mathematical Contest in Modeling (CUMCM), CSIAM. View certificate (PDF)
  • 2022.05, Finalist Award, Mathematical Contest in Modeling and Interdisciplinary Contest in Modeling (MCM/ICM), COMAP. View certificate (PDF)

💼 Research Experience

  • 2023.09 – present, Research Intern, Shanghai AI Laboratory.

🎓 Education

  • 2024.09 – present, Ph.D. student, Shanghai Jiao Tong University.

🤝 Academic Service

  • Reviewer for ICLR, NeurIPS (2025 Top Reviewer), CVPR, ICML, ECCV, and IEEE TMM.
  • Organizing committee member of VPLOW@CVPR 2024.