P
PaperPicks
Conferences
Mike Zheng Shou
National University of Singapore
69 papers at tracked venues · 59 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0000-0002-7681-2166 ↗
Homepage ↗
Venues
CVPR
×21
NeurIPS
×19
ECCV
×7
ICLR
×6
ACM MM
×4
ICCV
×4
AAAI
×2
ICML
×2
IJCAI
×2
ACL
×1
EMNLP
×1
Frequent coauthors
Yiren Song
DBLP profile ↗
ORCID search ↗
×4
Kevin Qinghong Lin
DBLP profile ↗
ORCID search ↗
×4
Zechen Bai
DBLP profile ↗
ORCID search ↗
×3
Rui Zhao
DBLP profile ↗
ORCID search ↗
×3
Henry Hengyuan Zhao
DBLP profile ↗
ORCID search ↗
×3
Yuchao Gu
DBLP profile ↗
ORCID search ↗
×3
Jinheng Xie
DBLP profile ↗
ORCID search ↗
×3
Ziteng Gao
DBLP profile ↗
ORCID search ↗
×3
Haiyang Mei
DBLP profile ↗
ORCID search ↗
×2
Xinyu Zhang
DBLP profile ↗
ORCID search ↗
×2
Weixian Lei
DBLP profile ↗
ORCID search ↗
×2
Joya Chen
DBLP profile ↗
ORCID search ↗
×2
Papers
OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization
AAAI 2026
·
Jiazheng Xing
DBLP profile ↗
ORCID search ↗
Balanced Image Stylization with Style Matching Score
ICCV 2025
·
Yuxin Jiang
DBLP profile ↗
ORCID search ↗
Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
ICLR 2025
·
Zechen Bai
DBLP profile ↗
ORCID search ↗
Can I Trust You? Advancing GUI Task Automation with Action Trust Score
ACM MM 2025
·
Haiyang Mei
DBLP profile ↗
ORCID search ↗
CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
NeurIPS 2025
·
Xinyu Zhang
DBLP profile ↗
ORCID search ↗
DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
CVPR 2025
·
Jay Zhangjie Wu
DBLP profile ↗
ORCID search ↗
DOTA: Distributional Test-time Adaptation of Vision-Language Models
NeurIPS 2025
·
Zongbo Han
DBLP profile ↗
ORCID search ↗
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity
ICCV 2025
·
Yiren Song
DBLP profile ↗
ORCID search ↗
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
CVPR 2025
·
Rui Zhao
DBLP profile ↗
ORCID search ↗
Factorized Learning for Temporally Grounded Video-Language Models
ICCV 2025
·
Wenzheng Zeng
DBLP profile ↗
ORCID search ↗
GUI-Narrator: Detecting and Captioning Computer GUI Actions
ACM MM 2025
·
Qinchen Wu
DBLP profile ↗
ORCID search ↗
Grounding Multimodal Large Language Model in GUI World
ICLR 2025
·
Weixian Lei
DBLP profile ↗
ORCID search ↗
IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation
CVPR 2025
·
Yiren Song
DBLP profile ↗
ORCID search ↗
Image Watermarks are Removable using Controllable Regeneration from Clean Noise
ICLR 2025
·
Yepeng Liu
DBLP profile ↗
ORCID search ↗
Impossible Videos
ICML 2025
·
Zechen Bai
DBLP profile ↗
ORCID search ↗
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models with Human Feedback
EMNLP 2025
·
Henry Hengyuan Zhao
DBLP profile ↗
ORCID search ↗
LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer
ICCV 2025
·
Yiren Song
DBLP profile ↗
ORCID search ↗
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
CVPR 2025
·
Joya Chen
DBLP profile ↗
ORCID search ↗
MP-Mat: A 3D-and-Instance-Aware Human Matting and Editing Framework with Multiplane Representation
ICLR 2025
·
Siyi Jiao
DBLP profile ↗
ORCID search ↗
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
CVPR 2025
·
Weijia Wu
DBLP profile ↗
ORCID search ↗
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
NeurIPS 2025
·
Yiren Song
DBLP profile ↗
ORCID search ↗
PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
NeurIPS 2025
·
Zhiwei Yang
DBLP profile ↗
ORCID search ↗
PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning
ACL 2025
·
Xinyu Zhang
DBLP profile ↗
ORCID search ↗
ROICtrl: Boosting Instance Control for Visual Generation
CVPR 2025
·
Yuchao Gu
DBLP profile ↗
ORCID search ↗
ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning
CVPR 2025
·
David Junhao Zhang
DBLP profile ↗
ORCID search ↗
SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
CVPR 2025
·
Haiyang Mei
DBLP profile ↗
ORCID search ↗
Show-o2: Improved Native Unified Multimodal Models
NeurIPS 2025
·
Jinheng Xie
DBLP profile ↗
ORCID search ↗
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
ICLR 2025
·
Jinheng Xie
DBLP profile ↗
ORCID search ↗
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
CVPR 2025
·
Kevin Qinghong Lin
DBLP profile ↗
ORCID search ↗
Sparse Image Synthesis via Joint Latent and RoI Flow
NeurIPS 2025
·
Ziteng Gao
DBLP profile ↗
ORCID search ↗
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
NeurIPS 2025
·
Jiaqi Wang
DBLP profile ↗
ORCID search ↗
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
AAAI 2025
·
Muhammet Furkan Ilaslan
DBLP profile ↗
ORCID search ↗
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
CVPR 2025
·
Kevin Qinghong Lin
DBLP profile ↗
ORCID search ↗
WMAdapter: Adding WaterMark Control to Latent Diffusion Models
ICML 2025
·
Hai Ci
DBLP profile ↗
ORCID search ↗
macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
NeurIPS 2025
·
Pei Yang
DBLP profile ↗
ORCID search ↗
Apprenticeship-Inspired Elegance: Synergistic Knowledge Distillation Empowers Spiking Neural Networks for Efficient Single-Eye Emotion Recognition
IJCAI 2024
·
Yang Wang
DBLP profile ↗
ORCID search ↗
AssistEditor: Multi-Agent Collaboration for GUI Workflow Automation in Video Creation
ACM MM 2024
·
Difei Gao
DBLP profile ↗
ORCID search ↗
AssistGUI: Task-Oriented PC Graphical User Interface Automation
CVPR 2024
·
Difei Gao
DBLP profile ↗
ORCID search ↗
Bootstrapping SparseFormers from Vision Foundation Models
CVPR 2024
·
Ziteng Gao
DBLP profile ↗
ORCID search ↗
Can Simple Averaging Defeat Modern Watermarks?
NeurIPS 2024
·
Pei Yang
DBLP profile ↗
ORCID search ↗
Delocate: Detection and Localization for Deepfake Videos with Randomly-Located Tampered Traces
IJCAI 2024
·
Juan Hu
DBLP profile ↗
ORCID search ↗
DoFIT: Domain-aware Federated Instruction Tuning with Alleviated Catastrophic Forgetting
NeurIPS 2024
·
Binqian Xu
DBLP profile ↗
ORCID search ↗
DragAnything: Motion Control for Anything Using Entity Representation
ECCV 2024
·
Weijia Wu
DBLP profile ↗
ORCID search ↗
DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing
CVPR 2024
·
Jia-Wei Liu
DBLP profile ↗
ORCID search ↗
EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models
NeurIPS 2024
·
Rui Zhao
DBLP profile ↗
ORCID search ↗
Exocentric-to-Egocentric Video Generation
NeurIPS 2024
·
Jia-Wei Liu
DBLP profile ↗
ORCID search ↗
Free-ATM: Harnessing Free Attention Masks for Representation Learning on Diffusion-Generated Images
ECCV 2024
·
David Junhao Zhang
DBLP profile ↗
ORCID search ↗
GENIXER: Empowering Multimodal Large Language Model as a Powerful Data Generator
ECCV 2024
·
Henry Hengyuan Zhao
DBLP profile ↗
ORCID search ↗
L4D-Track: Language-to-4D Modeling Towards 6-DoF Tracking and Shape Reconstruction in 3D Point Cloud Stream
CVPR 2024
·
Jingtao Sun
DBLP profile ↗
ORCID search ↗
LOVA3: Learning to Visual Question Answering, Asking and Assessment
NeurIPS 2024
·
Henry Hengyuan Zhao
DBLP profile ↗
ORCID search ↗
Learning Video Context as Interleaved Multimodal Sequences
ECCV 2024
·
Kevin Qinghong Lin
DBLP profile ↗
ORCID search ↗
Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
NeurIPS 2024
·
Alex Jinpeng Wang
DBLP profile ↗
ORCID search ↗
MAG-Edit: Localized Image Editing in Complex Scenarios via Mask-Based Attention-Adjusted Guidance
ACM MM 2024
·
Qi Mao
DBLP profile ↗
ORCID search ↗
MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model
CVPR 2024
·
Zhongcong Xu
DBLP profile ↗
ORCID search ↗
MotionDirector: Motion Customization of Text-to-Video Diffusion Models
ECCV 2024
·
Rui Zhao
DBLP profile ↗
ORCID search ↗
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
NeurIPS 2024
·
Zechen Bai
DBLP profile ↗
ORCID search ↗
Parrot Captions Teach CLIP to Spot Text
ECCV 2024
·
Yiqi Lin
DBLP profile ↗
ORCID search ↗
Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis
CVPR 2024
·
Yuchao Gu
DBLP profile ↗
ORCID search ↗
RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-key Identification
ECCV 2024
·
Hai Ci
DBLP profile ↗
ORCID search ↗
Skinned Motion Retargeting with Dense Geometric Interaction Perception
NeurIPS 2024
·
Zijie Ye
DBLP profile ↗
ORCID search ↗
SparseFormer: Sparse Visual Recognition via Limited Latent Tokens
ICLR 2024
·
Ziteng Gao
DBLP profile ↗
ORCID search ↗
Tune-an-Ellipse: CLIP Has Potential to Find what you Want
CVPR 2024
·
Jinheng Xie
DBLP profile ↗
ORCID search ↗
VIT-LENS: Towards Omni-modal Representations
CVPR 2024
·
Weixian Lei
DBLP profile ↗
ORCID search ↗
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
NeurIPS 2024
·
Kevin Qinghong Lin
DBLP profile ↗
ORCID search ↗
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
NeurIPS 2024
·
Shiwei Wu
DBLP profile ↗
ORCID search ↗
VideoLLM-online: Online Video Large Language Model for Streaming Video
CVPR 2024
·
Joya Chen
DBLP profile ↗
ORCID search ↗
VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point Correspondence
CVPR 2024
·
Yuchao Gu
DBLP profile ↗
ORCID search ↗
Visual Perception by Large Language Model's Weights
NeurIPS 2024
·
Feipeng Ma
DBLP profile ↗
ORCID search ↗
X- Adapter: Universal Compatibility of Plugins for Upgraded Diffusion Model
CVPR 2024
·
Lingmin Ran
DBLP profile ↗
ORCID search ↗