I am a Ph.D. student of Zhejiang University, supervised by Prof. Yueting Zhuang(庄越挺) and Prof. Siliang Tang(汤斯亮). I obtained my B.S. degree at Southeast University.
My research interest includes Vision-Language Understanding & Generation and Controllable Image/Video Generation & Editing. I have published 9 papers includes CVPR, ICCV, ICML, NeurIPS et.al.
I have developed a comprehensive multi-modal instruction editing dataset AnyEdit, which comprises 2.5 million high-quality editing pairs spanning 25 editing types for the community (2w+ downloads).
I have developed the first MLLM datasets hallucination detection framework HalluciDoctor and received extensive follow-up (93 cites) from the community.
I am expected to graduate in June 2026 and seeking job opportunities. Please feel free to contact me if you are interested!
🔥 News
- 2025.06: 🎉 DataTailor are accepted by ICCV2025! Codes are available!
- 2025.05: 🎉 Two papers are accepted by ICML 2025! OmniBench developed a multi-dimensional visual agent benchmark, and Similar proposed a Step-wise Multi-dimensional Generalist Reward Model.
- 2025.04: 🎉 AnyEdit is accepted by CVPR2025 as Oral presentation! We are delighted to release AnyEdit with code & dataset, which is a comprehensive multi-modal instruction editing dataset.
- 2025.02: 🎉 Two papers are accepted by CVPR 2025! AnyEdit received full scores (5,5,5) from the community
- 2024.09: 🎉 1 Paper is accepted by NeurIPS 2024.
- 2024.02: 🎉 HalluciDoctor is accepted by CVPR 2024! HalluciDoctor is the first to investigate the severe hallucination toxicity in existing MLLM datasets.
- 2023.08: 🔥 We released Baby-DALL3, which can annotate anything in visual tasks and generate anything just all in one pipeline with GPT-4.
- 2023.06: I received my first paper, CaCao, in ICCV 2023. Welcome to STAR and FORK!
📝 Publications
✍️ Controllable Image Generation

AnyEdit Unified High-Quality Image Edit with Any Idea
Qifan Yu, Wei Chow, Zhongqi Yue*, Kaihang Pan, Yang Wu, Xiaoyang Wan, Juncheng Li, Siliang Tang, Hanwang Zhang, Yueting Zhuang
- We present a comprehensive multi-modal image editing dataset, AnyEdit, to address the scarcity of high-quality instruction editing data for controllable image generation. This repository contains the official implementation, models, datasets, and data toolkit for the pipeline.
- AnyEdit comprises 2.5 million high-quality editing pairs spanning 25 editing types for the community, and achieves SOTA results on numerous editing benchmarks.
- Find out the official datasets and pre-trained checkpoint AnySD.
- We also developed and open-sourced an Benchmark to evaluate all baselines and our model more comprehensively.
-
PreprintInteractive data synthesis for systematic vision adaptation via llms-aigcs collaboration, Qifan Yu, Juncheng Li, Wentao Ye, Siliang Tang, Yueting Zhuang, Code -
Under ReviewDancing avatar: Pose and text-guided human motion videos synthesis with image diffusion model, Bosheng Qin, Wentao Ye, Qifan Yu, Siliang Tang, Yueting Zhuang -
Under ReviewSOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models, Haoyu Zheng, Qifan Yu, Binghe Yu, Yang Dai, Wenqiao Zhang, Juncheng Li, Siliang Tang, Yueting Zhuang
🙆 Vision-languag Understanding

HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
Qifan Yu, Juncheng Li, Longhui Wei, Liang Pang, Wentao Ye, Bosheng Qin, Siliang Tang, Qi Tian, Yueting Zhuang
HalluciDoctor is the first hallucination mitigating framework for the hallucinatory toxicity in MLLM datasets (LLaVA et al.).

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
Qifan Yu, Zhebei Shen, Zhongqi Yue*, Yang Wu, Wenqiao Zhang, Yunfei Li, Juncheng Li, Siliang Tang, Yueting Zhuang
DataTailor provides a principled and interpretable way for multi-modal data selection, enabling 85% training cost saving for SFT!
-
CVPR 2025STEP: Enhancing Video-LLMs’ Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training, Haiyi Qiu, Minghe Gao, Long Qian, Kaihang Pan, Qifan Yu, Juncheng Li, Wenjie Wang, Siliang Tang, Yueting Zhuang, Tat-Seng Chua -
ICCV 2023Visually-prompted language model for fine-grained scene graph generation in an open world, Qifan Yu, Juncheng Li, Yu Wu, Siliang Tang, Wei Ji, Yueting Zhuang
🤖 Visual Agent

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
Wendong Bu, Yang Wu, Qifan Yu*, Minghe Gao, Bingchen Miao, Zhenkui Zhang, Kaihang Pan, Yunfei Li, Mengze Li, Wei Ji, Juncheng Li, Siliang Tang, Yueting Zhuang
- We propose a novel self-generating, graph-based benchmark, OmniBench, for comprehensive agent evaluation at multiple granularities.
- OmniBench contains 36k graph-structured tasks across 20 scenarios, achieving a 91% human acceptance rate. We show a promising Demo to show its performance across various capabilities and paving the way for future advancements.
ICML 2025Boosting Virtual Agent Learning and Reasoning: A Step-wise, Multi-dimensional, and Generalist Reward Model with Benchmark, Bingchen Miao, Yang Wu, Minghe Gao, Qifan Yu, Wendong Bu, Wenqiao Zhang, Yunfei Li, Siliang Tang, Tat-Seng Chua, Juncheng Li
Others
-
NeurIPS 2024Unified Generative and Discriminative Training for Multi-modal Large Language Models, Wei Chow, Juncheng Li, Qifan Yu, Kaihang Pan, Hao Fei, Zhiqi Ge, Shuai Yang, Siliang Tang, Hanwang Zhang, Qianru Sun -
NeurIPS 2024 SpotlightTowards unified multimodal editing with enhanced knowledge collaboration , Kaihang Pan, Zhaoyu Fan, Juncheng Li, Qifan Yu, Hao Fei, Siliang Tang, Richang Hong, Hanwang Zhang, Qianru Sun
🎖 Honors and Awards
- 2023-2024 Award of Honor for Graduate
- 2022-2023 Honor for Graduates-Excellence in Academic (Moral education) Innovation
- 2019-2020 National Scholarship (Top 1%)
- 2018-2019 Principal’s Scholarship (Top 1%)
- 2025.05 ICME 2025 Inova Challenge First Place Award-Interleaved Text-Image Generation Track
📖 Educations
- 2021.09 - 2026.06, Ph.D., Zhejiang University, Hangzhou.
- 2017.09 - 2021.06, Undergraduate, Computer Science, Southeast University, Hangzhou.
- 2014.09 - 2017.06, Le Cheng Boarding School, Wenzhou, Zhejiang.
💻 Internships
- 2025.04 - now, Huawei Terminal BG, AI Engineer, Hangzhou.