I am a Ph.D. student at Shanghai Jiao Tong University (SJTU), advised by Prof. Dahua Lin. I conduct my research at Shanghai AI Laboratory, where I am jointly advised by Yuhang Zang and Jiaqi Wang.
My current research focuses on recursive self-improvement (RSI), agent harness engineering, and agentic post-training. I am interested in building agents that can continually improve through interaction, stronger training environments, and scalable feedback. I welcome discussions and collaborations in these areas.
📢 News
SEAgent is accepted by ICML 2026.
We release WildClawBench, a benchmark for real-world, long-horizon agent evaluation.
We release Visual-ERM and VisualCritic-RewardBench.
ARM-Thinker is accepted by CVPR 2026.
RAR is accepted by IEEE Transactions on Image Processing.
We release SPARK, a policy and reward co-evolving framework.
We release CODA, a dual-brain computer-use agent.
Earlier news
Visual-RFT is accepted by ICCV 2025.
InternLM-XComposer2.5-Reward is accepted by Findings of ACL 2025.
MIA-DPO is accepted by ICLR 2025.
MMDU and MMLongBench-Doc (Spotlight) are accepted by NeurIPS 2024.
🔬 Publications
🤖 Agents
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation Paper
An in-the-wild benchmark for evaluating long-horizon agents in realistic OpenClaw environments.
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience Paper
A computer-use agent that autonomously learns novel software through curriculum generation, experience, and reinforcement learning.
CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Paper
A trainable planner-executor framework for robust execution and generalization in scientific computing GUIs.
Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning First Author Paper
Reinforcement fine-tuning that equips open-source LVLMs with web search and visual coding capabilities.
🎯 RL & Reward Models
Visual-ERM: Reward Modeling for Visual Equivalence First Author Paper
A multimodal generative reward model that evaluates vision-to-code outputs in rendered visual space with fine-grained, interpretable feedback.
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning Paper
A multimodal reward model that reasons over visual inputs and actively uses tools to provide stronger feedback.
SPARK: Synergistic Policy And Reward Co-Evolving Framework First Author Paper
An on-policy framework in which policy and generative reward modeling improve together using recycled rollout supervision.
Visual-RFT: Visual Reinforcement Fine-Tuning First Author Paper
The first comprehensive adaptation of R1-style reinforcement fine-tuning to multimodal perception tasks with verifiable rewards.
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Paper
A general multimodal reward model for RL supervision, test-time response selection, and training-data filtering.
📚 Other Publications
MMLongBench-Doc: Benchmarking Long-Context Document Understanding with Visualizations Paper
A benchmark for understanding long documents containing text, tables, charts, images, and page layouts.
📝 Technical Reports
VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models Report
An open-source toolkit for reproducible evaluation of large multimodality models.
Intern-S2-Preview Model
A scientific multimodal foundation model with stronger reasoning and long-horizon agent capabilities.
🛠️ Open Source
Selected developer tools and Agent Skills. Research code is linked with each publication above.
ClaudeScope Project Lead Project
Parse and visualize Claude Code session trajectories, token usage, tool calls, timing, and failures.
gene.skill Project Lead Project
Extract, recombine, and mutate capability genes from existing Agent Skills to create new composite skills.
🏆 Honors and Awards
- 2025.09, Zhang Liangqi Scholarship, Shanghai Jiao Tong University.
- 2024.06, Excellent Bachelor’s Thesis and Outstanding Undergraduate Graduate.
- 2023.07, Meritorious Award, Mathematical Contest in Modeling and Interdisciplinary Contest in Modeling (MCM/ICM), COMAP. View certificate (PDF)
- 2022.05, National Second Prize, China Undergraduate Mathematical Contest in Modeling (CUMCM), CSIAM. View certificate (PDF)
- 2022.05, Finalist Award, Mathematical Contest in Modeling and Interdisciplinary Contest in Modeling (MCM/ICM), COMAP. View certificate (PDF)
💼 Research Experience
- 2023.09 – present, Research Intern, Shanghai AI Laboratory.
🎓 Education
- 2024.09 – present, Ph.D. student, Shanghai Jiao Tong University.
🤝 Academic Service
- Reviewer for ICLR, NeurIPS (2025 Top Reviewer), CVPR, ICML, ECCV, and IEEE TMM.
- Organizing committee member of VPLOW@CVPR 2024.