📝 Publications
✍️ Controllable Image Generation

AnyEdit Unified High-Quality Image Edit with Any Idea
Qifan Yu, Wei Chow, Zhongqi Yue*, Kaihang Pan, Yang Wu, Xiaoyang Wan, Juncheng Li, Siliang Tang, Hanwang Zhang, Yueting Zhuang
- We present a comprehensive multi-modal image editing dataset, AnyEdit, to address the scarcity of high-quality instruction editing data for controllable image generation. This repository contains the official implementation, models, datasets, and data toolkit for the pipeline.
- AnyEdit comprises 2.5 million high-quality editing pairs spanning 25 editing types for the community, and achieves SOTA results on numerous editing benchmarks.
- Find out the official datasets and pre-trained checkpoint AnySD.
- We also developed and open-sourced an Benchmark to evaluate all baselines and our model more comprehensively.
-
PreprintInteractive data synthesis for systematic vision adaptation via llms-aigcs collaboration, Qifan Yu, Juncheng Li, Wentao Ye, Siliang Tang, Yueting Zhuang, Code -
Under ReviewDancing avatar: Pose and text-guided human motion videos synthesis with image diffusion model, Bosheng Qin, Wentao Ye, Qifan Yu, Siliang Tang, Yueting Zhuang -
Under ReviewSOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models, Haoyu Zheng, Qifan Yu, Binghe Yu, Yang Dai, Wenqiao Zhang, Juncheng Li, Siliang Tang, Yueting Zhuang
🙆 Vision-languag Understanding

HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
Qifan Yu, Juncheng Li, Longhui Wei, Liang Pang, Wentao Ye, Bosheng Qin, Siliang Tang, Qi Tian, Yueting Zhuang
HalluciDoctor is the first hallucination mitigating framework for the hallucinatory toxicity in MLLM datasets (LLaVA et al.).

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
Qifan Yu, Zhebei Shen, Zhongqi Yue*, Yang Wu, Wenqiao Zhang, Yunfei Li, Juncheng Li, Siliang Tang, Yueting Zhuang
DataTailor provides a principled and interpretable way for multi-modal data selection, enabling 85% training cost saving for SFT!
-
CVPR 2025STEP: Enhancing Video-LLMs’ Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training, Haiyi Qiu, Minghe Gao, Long Qian, Kaihang Pan, Qifan Yu, Juncheng Li, Wenjie Wang, Siliang Tang, Yueting Zhuang, Tat-Seng Chua -
ICCV 2023Visually-prompted language model for fine-grained scene graph generation in an open world, Qifan Yu, Juncheng Li, Yu Wu, Siliang Tang, Wei Ji, Yueting Zhuang
🤖 Visual Agent

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
Wendong Bu, Yang Wu, Qifan Yu*, Minghe Gao, Bingchen Miao, Zhenkui Zhang, Kaihang Pan, Yunfei Li, Mengze Li, Wei Ji, Juncheng Li, Siliang Tang, Yueting Zhuang
- We propose a novel self-generating, graph-based benchmark, OmniBench, for comprehensive agent evaluation at multiple granularities.
- OmniBench contains 36k graph-structured tasks across 20 scenarios, achieving a 91% human acceptance rate. We show a promising Demo to show its performance across various capabilities and paving the way for future advancements.
ICML 2025Boosting Virtual Agent Learning and Reasoning: A Step-wise, Multi-dimensional, and Generalist Reward Model with Benchmark, Bingchen Miao, Yang Wu, Minghe Gao, Qifan Yu, Wendong Bu, Wenqiao Zhang, Yunfei Li, Siliang Tang, Tat-Seng Chua, Juncheng Li
Others
-
NeurIPS 2024Unified Generative and Discriminative Training for Multi-modal Large Language Models, Wei Chow, Juncheng Li, Qifan Yu, Kaihang Pan, Hao Fei, Zhiqi Ge, Shuai Yang, Siliang Tang, Hanwang Zhang, Qianru Sun -
NeurIPS 2024 SpotlightTowards unified multimodal editing with enhanced knowledge collaboration , Kaihang Pan, Zhaoyu Fan, Juncheng Li, Qifan Yu, Hao Fei, Siliang Tang, Richang Hong, Hanwang Zhang, Qianru Sun