P
PaperPicks
Conferences
Ming Yang
Ant Group, Seattle, WA, USA
27 papers at tracked venues · 23 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0000-0003-1691-6817 ↗
Google Scholar ↗
Homepage ↗
Venues
CVPR
×10
ACM MM
×3
ECCV
×3
AAAI
×2
ICCV
×2
NeurIPS
×2
ACL
×1
ICLR
×1
ICML
×1
IJCAI
×1
SIGIR
×1
Frequent coauthors
Zhanzhou Feng
DBLP profile ↗
ORCID search ↗
×2
Shuai Tan
DBLP profile ↗
ORCID search ↗
×2
Peiqi Chen
DBLP profile ↗
ORCID search ↗
×2
Tianqi Li
DBLP profile ↗
ORCID search ↗
×1
Yudong Han
DBLP profile ↗
ORCID search ↗
×1
Xiaolong Wang
DBLP profile ↗
ORCID search ↗
×1
Shuwei Shi
DBLP profile ↗
ORCID search ↗
×1
Haina Qin
DBLP profile ↗
ORCID search ↗
×1
Muzhi Zhu
DBLP profile ↗
ORCID search ↗
×1
Qi Zhu
DBLP profile ↗
ORCID search ↗
×1
Yuyan Chen
DBLP profile ↗
ORCID search ↗
×1
Zheng Qin
DBLP profile ↗
ORCID search ↗
×1
Papers
SCAN: Self-Calibrated AutoregressioN for High-Quality Visual Generation
AAAI 2026
·
Zhanzhou Feng
DBLP profile ↗
ORCID search ↗
Animate-X: Universal Character Image Animation with Enhanced Motion Representation
ICLR 2025
·
Shuai Tan
DBLP profile ↗
ORCID search ↗
CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for Guidance
ICCV 2025
·
Peiqi Chen
DBLP profile ↗
ORCID search ↗
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
ACM MM 2025
·
Tianqi Li
DBLP profile ↗
ORCID search ↗
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
CVPR 2025
·
Yudong Han
DBLP profile ↗
ORCID search ↗
HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography Estimation
AAAI 2025
·
Xiaolong Wang
DBLP profile ↗
ORCID search ↗
Mimir: Improving Video Diffusion Models for Precise Text Understanding
CVPR 2025
·
Shuai Tan
DBLP profile ↗
ORCID search ↗
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
CVPR 2025
·
Shuwei Shi
DBLP profile ↗
ORCID search ↗
Reversing Flow for Image Restoration
CVPR 2025
·
Haina Qin
DBLP profile ↗
ORCID search ↗
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
CVPR 2025
·
Muzhi Zhu
DBLP profile ↗
ORCID search ↗
SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling
CVPR 2025
·
Qi Zhu
DBLP profile ↗
ORCID search ↗
Unified Visual Generation via Next-Set Prediction in Continuous Domain
ICCV 2025
·
Zhanzhou Feng
DBLP profile ↗
ORCID search ↗
VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions
ACL 2025
·
Yuyan Chen
DBLP profile ↗
ORCID search ↗
Versatile Multimodal Controls for Expressive Talking Human Animation
ACM MM 2025
·
Zheng Qin
DBLP profile ↗
ORCID search ↗
Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
NeurIPS 2024
·
Ziyuan Huang
DBLP profile ↗
ORCID search ↗
EVE: Efficient Zero-Shot Text-Based Video Editing With Depth Map Guidance and Temporal Consistency Constraints
IJCAI 2024
·
Yutao Chen
DBLP profile ↗
ORCID search ↗
EcoMatcher: Efficient Clustering Oriented Matcher for Detector-Free Image Matching
ECCV 2024
·
Peiqi Chen
DBLP profile ↗
ORCID search ↗
Learning Dynamic Tetrahedra for High-Quality Talking Head Synthesis
CVPR 2024
·
Zicheng Zhang
DBLP profile ↗
ORCID search ↗
M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval
SIGIR 2024
·
Xingning Dong
DBLP profile ↗
ORCID search ↗
POA: Pre-training Once for Models of All Sizes
ECCV 2024
·
Yingying Zhang
DBLP profile ↗
ORCID search ↗
Parameter-Efficient Complementary Expert Learning for Long-Tailed Visual Recognition
ACM MM 2024
·
Lixiang Ru
DBLP profile ↗
ORCID search ↗
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
CVPR 2024
·
Shiyu Xuan
DBLP profile ↗
ORCID search ↗
Referencing Where to Focus: Improving Visual Grounding with Referential Query
NeurIPS 2024
·
Yabing Wang
DBLP profile ↗
ORCID search ↗
SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery
CVPR 2024
·
Xin Guo
DBLP profile ↗
ORCID search ↗
StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models
ECCV 2024
·
Wen Li
DBLP profile ↗
ORCID search ↗
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
ICML 2024
·
Ziping Ma
DBLP profile ↗
ORCID search ↗
Towards Better Vision-Inspired Vision-Language Models
CVPR 2024
·
Yun-Hao Cao
DBLP profile ↗
ORCID search ↗