P
PaperPicks
Conferences
Zhe Gan
20 papers at tracked venues · 15 at CORE A* · active 2024–2025
DBLP profile ↗
ORCID search ↗
Venues
ICLR
×8
ECCV
×4
CVPR
×3
AAAI
×1
ACL
×1
EMNLP
×1
ICCV
×1
ICML
×1
Frequent coauthors
Zhengfeng Lai
DBLP profile ↗
ORCID search ↗
×2
Tsu-Jui Fu
DBLP profile ↗
ORCID search ↗
×2
Hong-You Chen
DBLP profile ↗
ORCID search ↗
×1
Zhangheng Li
DBLP profile ↗
ORCID search ↗
×1
Andrew Szot
DBLP profile ↗
ORCID search ↗
×1
Ruohong Zhang
DBLP profile ↗
ORCID search ↗
×1
Yusu Qian
DBLP profile ↗
ORCID search ↗
×1
Haotian Zhang
DBLP profile ↗
ORCID search ↗
×1
Hanrong Ye
DBLP profile ↗
ORCID search ↗
×1
Enrico Fini
DBLP profile ↗
ORCID search ↗
×1
Dong Wang
DBLP profile ↗
ORCID search ↗
×1
Ajay Kumar Jaiswal
DBLP profile ↗
ORCID search ↗
×1
Papers
Contrastive Localized Language-Image Pre-Training
ICML 2025
·
Hong-You Chen
DBLP profile ↗
ORCID search ↗
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
ICLR 2025
·
Zhangheng Li
DBLP profile ↗
ORCID search ↗
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons
CVPR 2025
·
Andrew Szot
DBLP profile ↗
ORCID search ↗
Improve Vision Language Model Chain-of-thought Reasoning
ACL 2025
·
Ruohong Zhang
DBLP profile ↗
ORCID search ↗
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
ICLR 2025
·
Yusu Qian
DBLP profile ↗
ORCID search ↗
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
ICLR 2025
·
Haotian Zhang
DBLP profile ↗
ORCID search ↗
MMEgo: Towards Building Egocentric Multimodal LLMs for Video QA
ICLR 2025
·
Hanrong Ye
DBLP profile ↗
ORCID search ↗
Multimodal Autoregressive Pre-training of Large Vision Encoders
CVPR 2025
·
Enrico Fini
DBLP profile ↗
ORCID search ↗
Proactive Pseudo-Intervention: Pre-informed Contrastive Learning For Interpretable Vision Models
AAAI 2025
·
Dong Wang
DBLP profile ↗
ORCID search ↗
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
ICLR 2025
·
Zhengfeng Lai
DBLP profile ↗
ORCID search ↗
UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
ICCV 2025
·
Tsu-Jui Fu
DBLP profile ↗
ORCID search ↗
Compressing LLMs: The Truth is Rarely Pure and Never Simple
ICLR 2024
·
Ajay Kumar Jaiswal
DBLP profile ↗
ORCID search ↗
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
CVPR 2024
·
Jaemin Cho
DBLP profile ↗
ORCID search ↗
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
ECCV 2024
·
Keen You
DBLP profile ↗
ORCID search ↗
Ferret: Refer and Ground Anything Anywhere at Any Granularity
ICLR 2024
·
Haoxuan You
DBLP profile ↗
ORCID search ↗
GRiT: A Generative Region-to-Text Transformer for Object Understanding
ECCV 2024
·
Jialian Wu
DBLP profile ↗
ORCID search ↗
Guiding Instruction-based Image Editing via Multimodal Large Language Models
ICLR 2024
·
Tsu-Jui Fu
DBLP profile ↗
ORCID search ↗
MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training
ECCV 2024
·
Brandon McKinzie
DBLP profile ↗
ORCID search ↗
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
EMNLP 2024
·
Yuhui Zhang
DBLP profile ↗
ORCID search ↗
VeCLIP: Improving CLIP Training via Visual-Enriched Captions
ECCV 2024
·
Zhengfeng Lai
DBLP profile ↗
ORCID search ↗