P
PaperPicks
Conferences
Ying Shan
82 papers at tracked venues · 68 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0000-0001-7673-8325 ↗
Homepage ↗
Venues
CVPR
×28
ECCV
×12
ICCV
×11
AAAI
×9
NeurIPS
×7
ICLR
×5
ICML
×3
ACL
×2
ACM MM
×2
InterSpeech
×1
NAACL
×1
WWW
×1
Frequent coauthors
Chong Mou
DBLP profile ↗
ORCID search ↗
×4
Yicheng Xiao
DBLP profile ↗
ORCID search ↗
×3
Yi Chen
DBLP profile ↗
ORCID search ↗
×2
Tao Wu
DBLP profile ↗
ORCID search ↗
×2
Yuying Ge
DBLP profile ↗
ORCID search ↗
×2
Tian-Xing Xu
DBLP profile ↗
ORCID search ↗
×2
Xiangjun Gao
DBLP profile ↗
ORCID search ↗
×2
Chengyue Wu
DBLP profile ↗
ORCID search ↗
×2
Ye Liu
DBLP profile ↗
ORCID search ↗
×2
Zongyang Ma
DBLP profile ↗
ORCID search ↗
×2
Ruyang Liu
DBLP profile ↗
ORCID search ↗
×2
Xuan Ju
DBLP profile ↗
ORCID search ↗
×2
Papers
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
ACL 2026
·
Yi Chen
DBLP profile ↗
ORCID search ↗
MMhops-R1: Multimodal Multi-hop Reasoning
AAAI 2026
·
Tao Zhang
DBLP profile ↗
ORCID search ↗
AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
ICCV 2025
·
Junhao Cheng
DBLP profile ↗
ORCID search ↗
CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
AAAI 2025
·
Tao Wu
DBLP profile ↗
ORCID search ↗
DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creation
CVPR 2025
·
Wang Zhao
DBLP profile ↗
ORCID search ↗
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
CVPR 2025
·
Wenbo Hu
DBLP profile ↗
ORCID search ↗
DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation
ICCV 2025
·
Yuejiang Dong
DBLP profile ↗
ORCID search ↗
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
CVPR 2025
·
Minghong Cai
DBLP profile ↗
ORCID search ↗
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
CVPR 2025
·
Yuying Ge
DBLP profile ↗
ORCID search ↗
FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction
ICCV 2025
·
Jiale Xu
DBLP profile ↗
ORCID search ↗
GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers
ICCV 2025
·
Shijie Ma
DBLP profile ↗
ORCID search ↗
Geometrycrafter: Consistent Geometry Estimation for Open-World Videos With Diffusion Priors
ICCV 2025
·
Tian-Xing Xu
DBLP profile ↗
ORCID search ↗
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
ICML 2025
·
Rui Yang
DBLP profile ↗
ORCID search ↗
Image Conductor: Precision Control for Interactive Video Synthesis
AAAI 2025
·
Yaowei Li
DBLP profile ↗
ORCID search ↗
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
ICML 2025
·
Yicheng Xiao
DBLP profile ↗
ORCID search ↗
Mamba-3VL: Taming State Space Model for 3D Vision Language Learning
ICCV 2025
·
Yuan Wang
DBLP profile ↗
ORCID search ↗
Mani-GS: Gaussian Splatting Manipulation with Triangular Mesh
CVPR 2025
·
Xiangjun Gao
DBLP profile ↗
ORCID search ↗
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
NeurIPS 2025
·
Yicheng Xiao
DBLP profile ↗
ORCID search ↗
Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion
CVPR 2025
·
Songsong Yu
DBLP profile ↗
ORCID search ↗
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
ICCV 2025
·
Yi Chen
DBLP profile ↗
ORCID search ↗
NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images
CVPR 2025
·
Lingen Li
DBLP profile ↗
ORCID search ↗
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
NAACL 2025
·
Chengyue Wu
DBLP profile ↗
ORCID search ↗
SEED-Story: Multimodal Long Story Generation with Large Language Model
ICCV 2025
·
Shuai Yang
DBLP profile ↗
ORCID search ↗
Scalable Image Tokenization with Index Backpropagation Quantization
ICCV 2025
·
Fengyuan Shi
DBLP profile ↗
ORCID search ↗
Taming Rectified Flow for Inversion and Editing
ICML 2025
·
Jiangshan Wang
DBLP profile ↗
ORCID search ↗
TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
ICCV 2025
·
Mark Yu
DBLP profile ↗
ORCID search ↗
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
NeurIPS 2025
·
Ye Liu
DBLP profile ↗
ORCID search ↗
VisionMath: Vision-Form Mathematical Problem-Solving
ICCV 2025
·
Zongyang Ma
DBLP profile ↗
ORCID search ↗
A Pre-convolved Representation for Plug-and-Play Neural Illumination Fields
AAAI 2024
·
Yiyu Zhuang
DBLP profile ↗
ORCID search ↗
AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
InterSpeech 2024
·
Yongkang Yin
DBLP profile ↗
ORCID search ↗
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
CVPR 2024
·
Ruyang Liu
DBLP profile ↗
ORCID search ↗
BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion
ECCV 2024
·
Xuan Ju
DBLP profile ↗
ORCID search ↗
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
NeurIPS 2024
·
Sijie Zhao
DBLP profile ↗
ORCID search ↗
ConTex-Human: Free-View Rendering of Human from a Single Image with Texture-Consistent Synthesis
CVPR 2024
·
Xiangjun Gao
DBLP profile ↗
ORCID search ↗
CustomNet: Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models
ACM MM 2024
·
Ziyang Yuan
DBLP profile ↗
ORCID search ↗
DMiT: Deformable Mipmapped Tri-Plane Representation for Dynamic Scenes
ECCV 2024
·
Jing-Wen Yang
DBLP profile ↗
ORCID search ↗
DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing
CVPR 2024
·
Chong Mou
DBLP profile ↗
ORCID search ↗
DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models
ICLR 2024
·
Chong Mou
DBLP profile ↗
ORCID search ↗
DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion Models
CVPR 2024
·
Yukang Cao
DBLP profile ↗
ORCID search ↗
DreamDiffusion: High-Quality EEG-to-Image Generation with Temporal Masked Signal Modeling and CLIP Alignment
ECCV 2024
·
Yunpeng Bai
DBLP profile ↗
ORCID search ↗
DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing
CVPR 2024
·
Jia-Wei Liu
DBLP profile ↗
ORCID search ↗
DynamiCrafter: Animating Open-Domain Images with Video Diffusion Priors
ECCV 2024
·
Jinbo Xing
DBLP profile ↗
ORCID search ↗
E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
NeurIPS 2024
·
Ye Liu
DBLP profile ↗
ORCID search ↗
EA-VTR: Event-Aware Video-Text Retrieval
ECCV 2024
·
Zongyang Ma
DBLP profile ↗
ORCID search ↗
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
CVPR 2024
·
Yaofang Liu
DBLP profile ↗
ORCID search ↗
FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling
ICLR 2024
·
Haonan Qiu
DBLP profile ↗
ORCID search ↗
GS-IR: 3D Gaussian Splatting for Inverse Rendering
CVPR 2024
·
Zhihao Liang
DBLP profile ↗
ORCID search ↗
HiFi-123: Towards High-Fidelity One Image to 3D Content Generation
ECCV 2024
·
Wangbo Yu
DBLP profile ↗
ORCID search ↗
How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?
CVPR 2024
·
Yuxin Chen
DBLP profile ↗
ORCID search ↗
HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting
CVPR 2024
·
Xian Liu
DBLP profile ↗
ORCID search ↗
HumanRef: Single Image to 3D Human Generation via Reference-Guided Diffusion
CVPR 2024
·
Jingbo Zhang
DBLP profile ↗
ORCID search ↗
LLaMA Pro: Progressive LLaMA with Block Expansion
ACL 2024
·
Chengyue Wu
DBLP profile ↗
ORCID search ↗
Low-Rank Approximation for Sparse Attention in Multi-Modal LLMs
CVPR 2024
·
Lin Song
DBLP profile ↗
ORCID search ↗
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
ECCV 2024
·
Muyao Niu
DBLP profile ↗
ORCID search ↗
Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation
ECCV 2024
·
Lanqing Guo
DBLP profile ↗
ORCID search ↗
Making LLaMA SEE and Draw with SEED Tokenizer
ICLR 2024
·
Yuying Ge
DBLP profile ↗
ORCID search ↗
MambaTree: Tree Topology is All You Need in State Space Model
NeurIPS 2024
·
Yicheng Xiao
DBLP profile ↗
ORCID search ↗
MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
NeurIPS 2024
·
Xuan Ju
DBLP profile ↗
ORCID search ↗
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
CVPR 2024
·
Yiyuan Zhang
DBLP profile ↗
ORCID search ↗
Noise Calibration: Plug-and-Play Content-Preserving Video Enhancement Using Pre-trained Video Diffusion Models
ECCV 2024
·
Qinyu Yang
DBLP profile ↗
ORCID search ↗
PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
CVPR 2024
·
Zhen Li
DBLP profile ↗
ORCID search ↗
Programmable Motion Generation for Open-Set Motion Control Tasks
CVPR 2024
·
Hanchao Liu
DBLP profile ↗
ORCID search ↗
ReVideo: Remake a Video with Motion and Content Control
NeurIPS 2024
·
Chong Mou
DBLP profile ↗
ORCID search ↗
RecDCL: Dual Contrastive Learning for Recommendation
WWW 2024
·
Dan Zhang
DBLP profile ↗
ORCID search ↗
Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis
CVPR 2024
·
Yuchao Gu
DBLP profile ↗
ORCID search ↗
SC-NeuS: Consistent Neural Surface Reconstruction from Sparse and Noisy Views
AAAI 2024
·
Shi-Sheng Huang
DBLP profile ↗
ORCID search ↗
SEED-Bench: Benchmarking Multimodal Large Language Models
CVPR 2024
·
Bohao Li
DBLP profile ↗
ORCID search ↗
ST-LLM: Large Language Models Are Effective Temporal Learners
ECCV 2024
·
Ruyang Liu
DBLP profile ↗
ORCID search ↗
ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion Models
ICLR 2024
·
Yingqing He
DBLP profile ↗
ORCID search ↗
SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models
CVPR 2024
·
Yuzhou Huang
DBLP profile ↗
ORCID search ↗
Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse Views
AAAI 2024
·
Zixin Zou
DBLP profile ↗
ORCID search ↗
SparseGNV: Generating Novel Views of Indoor Scenes with Sparse RGB-D Images
AAAI 2024
·
Weihao Cheng
DBLP profile ↗
ORCID search ↗
SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model
AAAI 2024
·
Tao Wu
DBLP profile ↗
ORCID search ↗
Storytelling Video Generation with Retrieval Augmentation and Character Consistency
ECCV 2024
·
Yingqing He
DBLP profile ↗
ORCID search ↗
SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
ACM MM 2024
·
Chaolei Tan
DBLP profile ↗
ORCID search ↗
T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion Models
AAAI 2024
·
Chong Mou
DBLP profile ↗
ORCID search ↗
TapMo: Shape-aware Motion Generation of Skeleton-free Characters
ICLR 2024
·
Jiaxu Zhang
DBLP profile ↗
ORCID search ↗
Texture-GS: Disentangling the Geometry and Texture for 3D Gaussian Splatting Editing
ECCV 2024
·
Tian-Xing Xu
DBLP profile ↗
ORCID search ↗
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
CVPR 2024
·
Xiaohan Ding
DBLP profile ↗
ORCID search ↗
VIT-LENS: Towards Omni-modal Representations
CVPR 2024
·
Weixian Lei
DBLP profile ↗
ORCID search ↗
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
CVPR 2024
·
Haoxin Chen
DBLP profile ↗
ORCID search ↗
YOLO-World: Real-Time Open-Vocabulary Object Detection
CVPR 2024
·
Tianheng Cheng
DBLP profile ↗
ORCID search ↗