P
PaperPicks
Conferences
Di Zhang
Kuaishou Technology, Beijing, China
58 papers at tracked venues · 50 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0009-0006-5475-2728 ↗
Homepage ↗
Venues
CVPR
×10
ICCV
×8
NeurIPS
×8
ACL
×7
ICLR
×7
EMNLP
×6
AAAI
×4
ICML
×3
ACM MM
×2
NAACL
×2
SIGKDD
×1
Frequent coauthors
Junxian Li
DBLP profile ↗
ORCID search ↗
×2
Jie Liu
DBLP profile ↗
ORCID search ↗
×2
Longrong Yang
DBLP profile ↗
ORCID search ↗
×2
Zhicheng Zhang
DBLP profile ↗
ORCID search ↗
×2
Jianhong Bai
DBLP profile ↗
ORCID search ↗
×2
Jiao Ou
DBLP profile ↗
ORCID search ↗
×2
Yang Jin
DBLP profile ↗
ORCID search ↗
×2
Liang Hou
DBLP profile ↗
ORCID search ↗
×1
Xiangyang Luo
DBLP profile ↗
ORCID search ↗
×1
Yunxiao Wang
DBLP profile ↗
ORCID search ↗
×1
Xiao Fu
DBLP profile ↗
ORCID search ↗
×1
Haonan He
DBLP profile ↗
ORCID search ↗
×1
Papers
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
AAAI 2026
·
Liang Hou
DBLP profile ↗
ORCID search ↗
Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
ACL 2026
·
Junxian Li
DBLP profile ↗
ORCID search ↗
FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
AAAI 2026
·
Xiangyang Luo
DBLP profile ↗
ORCID search ↗
TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs
AAAI 2026
·
Yunxiao Wang
DBLP profile ↗
ORCID search ↗
3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation
ICLR 2025
·
Xiao Fu
DBLP profile ↗
ORCID search ↗
Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models
EMNLP 2025
·
Haonan He
DBLP profile ↗
ORCID search ↗
Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
ICLR 2025
·
Hejia Chen
DBLP profile ↗
ORCID search ↗
ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
AAAI 2025
·
Junxian Li
DBLP profile ↗
ORCID search ↗
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
CVPR 2025
·
Di Zhang
DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMs
EMNLP 2025
·
Minxuan Lv
DBLP profile ↗
ORCID search ↗
Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models
NeurIPS 2025
·
Wei Chen
DBLP profile ↗
ORCID search ↗
Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
NeurIPS 2025
·
Tao Zhang
DBLP profile ↗
ORCID search ↗
Flow-GRPO: Training Flow Matching Models via Online RL
NeurIPS 2025
·
Jie Liu
DBLP profile ↗
ORCID search ↗
FullDiT: Video Generative Foundation Models with Multimodal Control via Full Attention
ICCV 2025
·
Xuan Ju
DBLP profile ↗
ORCID search ↗
GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation
ICCV 2025
·
Wentao Hu
DBLP profile ↗
ORCID search ↗
GPAvatar: High-fidelity Head Avatars by Learning Efficient Gaussian Projections
CVPR 2025
·
Wei-Qi Feng
DBLP profile ↗
ORCID search ↗
GameFactorly: Creating New Games with Generative Interactive Videos
ICCV 2025
·
Jiwen Yu
DBLP profile ↗
ORCID search ↗
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
ACL 2025
·
Xiao Wang
DBLP profile ↗
ORCID search ↗
How Far are AI-Generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach
ICCV 2025
·
Chirui Chang
DBLP profile ↗
ORCID search ↗
Imbalance in Balance: Online Concept Balancing in Generation Models
ICCV 2025
·
Yukai Shi
DBLP profile ↗
ORCID search ↗
Improving Video Generation with Human Feedback
NeurIPS 2025
·
Jie Liu
DBLP profile ↗
ORCID search ↗
Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
CVPR 2025
·
Qiuheng Wang
DBLP profile ↗
ORCID search ↗
LLaMA-Berry: Pairwise Optimization for Olympiad-level Mathematical Reasoning via O1-like Monte Carlo Tree Search
NAACL 2025
·
Di Zhang
Libra-Merging: Importance-redundancy and Pruning-merging Trade-off for Acceleration Plug-in in Large Vision-Language Model
CVPR 2025
·
Longrong Yang
DBLP profile ↗
ORCID search ↗
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
ICML 2025
·
Yifan Zhang
DBLP profile ↗
ORCID search ↗
MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
ICML 2025
·
Zhicheng Zhang
DBLP profile ↗
ORCID search ↗
MUSE: Multi-Subject Unified Synthesis Via Explicit Layout Semantic Expansion
ICCV 2025
·
Fei Peng
DBLP profile ↗
ORCID search ↗
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
ACM MM 2025
·
Yang Shi
DBLP profile ↗
ORCID search ↗
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
NeurIPS 2025
·
Ziqiao Peng
DBLP profile ↗
ORCID search ↗
PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
CVPR 2025
·
Shian Du
DBLP profile ↗
ORCID search ↗
Recammaster: Camera-Controlled Generative Rendering From a Single Video
ICCV 2025
·
Jianhong Bai
DBLP profile ↗
ORCID search ↗
Retrieval is Not Enough: Enhancing RAG through Test-Time Critique and Optimization
NeurIPS 2025
·
Jiaqi Wei
DBLP profile ↗
ORCID search ↗
SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin
EMNLP 2025
·
Hao Yi
DBLP profile ↗
ORCID search ↗
Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural Rectification
ICCV 2025
·
Guibao Shen
DBLP profile ↗
ORCID search ↗
SketchVideo: Sketch-based Video Generation and Editing
CVPR 2025
·
Feng-Lin Liu
DBLP profile ↗
ORCID search ↗
Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model
ICLR 2025
·
Longrong Yang
DBLP profile ↗
ORCID search ↗
Stable Segment Anything Model
ICLR 2025
·
Qi Fan
DBLP profile ↗
ORCID search ↗
StyleMaster: Stylize Your Video with Artistic Generation and Translation
CVPR 2025
·
Zixuan Ye
DBLP profile ↗
ORCID search ↗
SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
ICLR 2025
·
Jianhong Bai
DBLP profile ↗
ORCID search ↗
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
ICLR 2025
·
Jiankang Chen
DBLP profile ↗
ORCID search ↗
Towards Precise Scaling Laws for Video Diffusion Transformers
CVPR 2025
·
Yuanyang Yin
DBLP profile ↗
ORCID search ↗
Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation
CVPR 2025
·
Zhuoman Liu
DBLP profile ↗
ORCID search ↗
VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform
SIGKDD 2025
·
Xingyu Lu
DBLP profile ↗
ORCID search ↗
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
ACL 2025
·
Xinlong Chen
DBLP profile ↗
ORCID search ↗
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
NeurIPS 2025
·
Zhicheng Zhang
DBLP profile ↗
ORCID search ↗
iMOVE : Instance-Motion-Aware Video Understanding
ACL 2025
·
Jiaze Li
DBLP profile ↗
ORCID search ↗
DialogBench: Evaluating LLMs as Human-like Dialogue Systems
NAACL 2024
·
Jiao Ou
DBLP profile ↗
ORCID search ↗
Evaluating Readability and Faithfulness of Concept-based Explanations
EMNLP 2024
·
Meng Li
DBLP profile ↗
ORCID search ↗
Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint
ACL 2024
·
Zhipeng Chen
DBLP profile ↗
ORCID search ↗
Inductive-Deductive Strategy Reuse for Multi-Turn Instructional Dialogues
EMNLP 2024
·
Jiao Ou
DBLP profile ↗
ORCID search ↗
Just Ask One More Time! Self-Agreement Improves Reasoning of Language Models in (Almost) All Scenarios
ACL 2024
·
Lei Lin
DBLP profile ↗
ORCID search ↗
Learning Multi-Dimensional Human Preference for Text-to-Image Generation
CVPR 2024
·
Sixian Zhang
DBLP profile ↗
ORCID search ↗
Parrot: Enhancing Multi-Turn Instruction Following for Large Language Models
ACL 2024
·
Yuchong Sun
DBLP profile ↗
ORCID search ↗
PlacidDreamer: Advancing Harmony in Text-to-3D Generation
ACM MM 2024
·
Shuo Huang
DBLP profile ↗
ORCID search ↗
Small Agent Can Also Rock! Empowering Small Language Models as Hallucination Detector
EMNLP 2024
·
Xiaoxue Cheng
DBLP profile ↗
ORCID search ↗
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
ICLR 2024
·
Yang Jin
DBLP profile ↗
ORCID search ↗
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
ICML 2024
·
Yang Jin
DBLP profile ↗
ORCID search ↗
VideoTetris: Towards Compositional Text-to-Video Generation
NeurIPS 2024
·
Ye Tian
DBLP profile ↗
ORCID search ↗