📝 Publications

✍️ Controllable Image Generation

CVPR 2025 Oral
sym

AnyEdit Unified High-Quality Image Edit with Any Idea
Qifan Yu, Wei Chow, Zhongqi Yue*, Kaihang Pan, Yang Wu, Xiaoyang Wan, Juncheng Li, Siliang Tang, Hanwang Zhang, Yueting Zhuang

Project

  • We present a comprehensive multi-modal image editing dataset, AnyEdit, to address the scarcity of high-quality instruction editing data for controllable image generation. This repository contains the official implementation, models, datasets, and data toolkit for the pipeline.
  • AnyEdit comprises 2.5 million high-quality editing pairs spanning 25 editing types for the community, and achieves SOTA results on numerous editing benchmarks.
  • Find out the official datasets and pre-trained checkpoint AnySD.
  • We also developed and open-sourced an Benchmark to evaluate all baselines and our model more comprehensively.

🙆 Vision-languag Understanding

CVPR 2024
sym

HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
Qifan Yu, Juncheng Li, Longhui Wei, Liang Pang, Wentao Ye, Bosheng Qin, Siliang Tang, Qi Tian, Yueting Zhuang

Project

HalluciDoctor is the first hallucination mitigating framework for the hallucinatory toxicity in MLLM datasets (LLaVA et al.).

ICCV 2025
sym

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
Qifan Yu, Zhebei Shen, Zhongqi Yue*, Yang Wu, Wenqiao Zhang, Yunfei Li, Juncheng Li, Siliang Tang, Yueting Zhuang

Project

DataTailor provides a principled and interpretable way for multi-modal data selection, enabling 85% training cost saving for SFT!

🤖 Visual Agent

ICML 2025 Oral
sym

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
Wendong Bu, Yang Wu, Qifan Yu*, Minghe Gao, Bingchen Miao, Zhenkui Zhang, Kaihang Pan, Yunfei Li, Mengze Li, Wei Ji, Juncheng Li, Siliang Tang, Yueting Zhuang

Project

  • We propose a novel self-generating, graph-based benchmark, OmniBench, for comprehensive agent evaluation at multiple granularities.
  • OmniBench contains 36k graph-structured tasks across 20 scenarios, achieving a 91% human acceptance rate. We show a promising Demo to show its performance across various capabilities and paving the way for future advancements.

Others