P
PaperPicks
Conferences
Jingdong Wang
Baidu, AI Group, Sunnyvale, CA, USA
56 papers at tracked venues · 43 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0000-0002-4888-4445 ↗
Google Scholar ↗
Homepage ↗
Venues
CVPR
×16
NeurIPS
×12
ECCV
×10
AAAI
×7
ICLR
×3
ICML
×3
ACM MM
×2
ICRA
×2
IJCAI
×1
Frequent coauthors
Hao Li
DBLP profile ↗
ORCID search ↗
×3
Mengmeng Wang
DBLP profile ↗
ORCID search ↗
×2
Jiazhi Guan
DBLP profile ↗
ORCID search ↗
×2
Ziru Wang
DBLP profile ↗
ORCID search ↗
×2
Jiahao Cui
DBLP profile ↗
ORCID search ↗
×2
Yuzhe Yao
DBLP profile ↗
ORCID search ↗
×2
Haonan Lin
DBLP profile ↗
ORCID search ↗
×2
Chengyou Jia
DBLP profile ↗
ORCID search ↗
×2
Zhe Liu
DBLP profile ↗
ORCID search ↗
×2
Chuyang Zhao
DBLP profile ↗
ORCID search ↗
×2
Ze Feng
DBLP profile ↗
ORCID search ↗
×1
Zebin You
DBLP profile ↗
ORCID search ↗
×1
Papers
EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
AAAI 2026
·
Ze Feng
DBLP profile ↗
ORCID search ↗
Action Detail Matters: Refining Video Recognition with Local Action Queries
CVPR 2025
·
Mengmeng Wang
DBLP profile ↗
ORCID search ↗
Are Images Indistinguishable to Humans Also Indistinguishable to Classifiers?
CVPR 2025
·
Zebin You
DBLP profile ↗
ORCID search ↗
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
CVPR 2025
·
Jiazhi Guan
DBLP profile ↗
ORCID search ↗
Continual SFT Matches Multimodal RLHF with Negative Supervision
CVPR 2025
·
Ke Zhu
DBLP profile ↗
ORCID search ↗
DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes
ICRA 2025
·
Hao Li
DBLP profile ↗
ORCID search ↗
DynaMind: Reasoning over Abstract Video Dynamics for Embodied Decision-Making
ICML 2025
·
Ziru Wang
DBLP profile ↗
ORCID search ↗
Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
ICLR 2025
·
Jiahao Cui
DBLP profile ↗
ORCID search ↗
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
CVPR 2025
·
Jiahao Cui
DBLP profile ↗
ORCID search ↗
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
AAAI 2025
·
Guosheng Zhang
DBLP profile ↗
ORCID search ↗
Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving
ICRA 2025
·
Lingyu Xiao
DBLP profile ↗
ORCID search ↗
MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction
ICLR 2025
·
Jing Yang
DBLP profile ↗
ORCID search ↗
Manifold Constraint Reduces Exposure Bias in Accelerated Diffusion Sampling
ICLR 2025
·
Yuzhe Yao
DBLP profile ↗
ORCID search ↗
MonoLift: Learning 3D Manipulation Policies from Monocular RGB via Distillation
NeurIPS 2025
·
Ziru Wang
DBLP profile ↗
ORCID search ↗
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
CVPR 2025
·
Hui Li
DBLP profile ↗
ORCID search ↗
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
CVPR 2025
·
Yingying Fan
DBLP profile ↗
ORCID search ↗
SpotActor: Training-Free Layout-Controlled Consistent Image Generation
AAAI 2025
·
Jiahao Wang
DBLP profile ↗
ORCID search ↗
TexGarment: Consistent Garment UV Texture Generation via Efficient 3D Structure-Guided Diffusion Transformer
CVPR 2025
·
Jialun Liu
DBLP profile ↗
ORCID search ↗
TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP
ACM MM 2025
·
Fan Li
DBLP profile ↗
ORCID search ↗
VidEvo: Evolving Video Editing through Exhaustive Temporal Modeling
IJCAI 2025
·
Sizhe Dang
DBLP profile ↗
ORCID search ↗
VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction
CVPR 2025
·
Ziyue Zhu
DBLP profile ↗
ORCID search ↗
3D-Aware Text-Driven Talking Avatar Generation
ECCV 2024
·
Xiuzhe Wu
DBLP profile ↗
ORCID search ↗
A Multimodal, Multi-Task Adapting Framework for Video Action Recognition
AAAI 2024
·
Mengmeng Wang
DBLP profile ↗
ORCID search ↗
Automated Multi-level Preference for MLLMs
NeurIPS 2024
·
Mengxi Zhang
DBLP profile ↗
ORCID search ↗
BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object Detection
CVPR 2024
·
Wenjie Wang
DBLP profile ↗
ORCID search ↗
Decoupled Pseudo-Labeling for Semi-Supervised Monocular 3D Object Detection
CVPR 2024
·
Jiacheng Zhang
DBLP profile ↗
ORCID search ↗
Dense Connector for MLLMs
NeurIPS 2024
·
Huanjin Yao
DBLP profile ↗
ORCID search ↗
Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
NeurIPS 2024
·
Mingxiang Liao
DBLP profile ↗
ORCID search ↗
Flipped Classroom: Aligning Teacher Attention with Student in Generalized Category Discovery
NeurIPS 2024
·
Haonan Lin
DBLP profile ↗
ORCID search ↗
Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection
CVPR 2024
·
Huan Liu
DBLP profile ↗
ORCID search ↗
GGRt: Towards Pose-Free Generalizable 3D Gaussian Splatting in Real-Time
ECCV 2024
·
Hao Li
DBLP profile ↗
ORCID search ↗
GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding
CVPR 2024
·
Hao Li
DBLP profile ↗
ORCID search ↗
Generating Action-conditioned Prompts for Open-vocabulary Video Action Recognition
ACM MM 2024
·
Chengyou Jia
DBLP profile ↗
ORCID search ↗
IRGen: Generative Modeling for Image Retrieval
ECCV 2024
·
Yidan Zhang
DBLP profile ↗
ORCID search ↗
Interactive 3D Object Detection with Prompts
ECCV 2024
·
Rui Zhang
DBLP profile ↗
ORCID search ↗
LION: Linear Group RNN for 3D Object Detection in Point Clouds
NeurIPS 2024
·
Zhe Liu
DBLP profile ↗
ORCID search ↗
LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction
ECCV 2024
·
Penghui Du
DBLP profile ↗
ORCID search ↗
Learning to Rematch Mismatched Pairs for Robust Cross-Modal Retrieval
CVPR 2024
·
Haochen Han
DBLP profile ↗
ORCID search ↗
MS-DETR: Efficient DETR Training with Mixed Supervision
CVPR 2024
·
Chuyang Zhao
DBLP profile ↗
ORCID search ↗
Make Your ViT-Based Multi-view 3D Detectors Faster via Token Compression
ECCV 2024
·
Dingyuan Zhang
DBLP profile ↗
ORCID search ↗
MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
NeurIPS 2024
·
Jie Zhu
DBLP profile ↗
ORCID search ↗
Mobile Attention: Mobile-Friendly Linear-Attention for Vision Transformers
ICML 2024
·
Zhiyu Yao
DBLP profile ↗
ORCID search ↗
Multi-Domain Incremental Learning for Face Presentation Attack Detection
AAAI 2024
·
Keyao Wang
DBLP profile ↗
ORCID search ↗
Noisy Correspondence Learning with Self-Reinforcing Errors Mitigation
AAAI 2024
·
Zhuohang Dang
DBLP profile ↗
ORCID search ↗
OPEN: Object-Wise Position Embedding for Multi-view 3D Object Detection
ECCV 2024
·
Jinghua Hou
DBLP profile ↗
ORCID search ↗
Octopus: A Multi-modal LLM with Parallel Recognition and Sequential Understanding
NeurIPS 2024
·
Chuyang Zhao
DBLP profile ↗
ORCID search ↗
OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding
NeurIPS 2024
·
Yanmin Wu
DBLP profile ↗
ORCID search ↗
PLIP: Language-Image Pre-training for Person Representation Learning
NeurIPS 2024
·
Jialong Zuo
DBLP profile ↗
ORCID search ↗
ReSyncer: Rewiring Style-Based Generator for Unified Audio-Visually Synced Facial Performer
ECCV 2024
·
Jiazhi Guan
DBLP profile ↗
ORCID search ↗
SEED: A Simple and Effective 3D DETR in Point Clouds
ECCV 2024
·
Zhe Liu
DBLP profile ↗
ORCID search ↗
SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-Form Layout-to-Image Generation
AAAI 2024
·
Chengyou Jia
DBLP profile ↗
ORCID search ↗
Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image Editing
NeurIPS 2024
·
Haonan Lin
DBLP profile ↗
ORCID search ↗
ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion Modeling
NeurIPS 2024
·
Quanwei Yang
DBLP profile ↗
ORCID search ↗
Timestep-Aware Correction for Quantized Diffusion Models
ECCV 2024
·
Yuzhe Yao
DBLP profile ↗
ORCID search ↗
Towards Unified Multi-granularity Text Detection with Interactive Attention
ICML 2024
·
Xingyu Wan
DBLP profile ↗
ORCID search ↗
VRP-SAM: SAM with Visual Reference Prompt
CVPR 2024
·
Yanpeng Sun
DBLP profile ↗
ORCID search ↗