P
PaperPicks
Conferences
Xing Sun
Tencent Youtu Lab, Shanghai, China
36 papers at tracked venues · 33 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0000-0001-8132-9083 ↗
Google Scholar ↗
Homepage ↗
Venues
ACL
×7
NeurIPS
×7
AAAI
×5
CVPR
×5
ACM MM
×4
ICML
×3
EMNLP
×2
ICLR
×2
ECCV
×1
Frequent coauthors
Chaoyou Fu
DBLP profile ↗
ORCID search ↗
×3
Chuang Zhou
DBLP profile ↗
ORCID search ↗
×2
Xin Li
DBLP profile ↗
ORCID search ↗
×2
Zhen Sun
DBLP profile ↗
ORCID search ↗
×2
Junru Lu
DBLP profile ↗
ORCID search ↗
×2
Juntao Wu
DBLP profile ↗
ORCID search ↗
×1
Wensheng Lu
DBLP profile ↗
ORCID search ↗
×1
Xiong Wang
DBLP profile ↗
ORCID search ↗
×1
Yulei Qin
DBLP profile ↗
ORCID search ↗
×1
Liuhao Lin
DBLP profile ↗
ORCID search ↗
×1
Chenyu Zhou
DBLP profile ↗
ORCID search ↗
×1
Shuyang Liu
DBLP profile ↗
ORCID search ↗
×1
Papers
Breaking the Evaluation Paradox: Evaluating High-Entropy Search with Computationally Irreducible Constraints
ACL 2026
·
Juntao Wu
DBLP profile ↗
ORCID search ↗
Collision to Cognition: Hash-Driven Graph Construction for Efficient RAG
ACL 2026
·
Chuang Zhou
DBLP profile ↗
ORCID search ↗
HiChunk: Evaluating and Enhancing Retrieval Augmented Generation with Hierarchical Chunking
ACL 2026
·
Wensheng Lu
DBLP profile ↗
ORCID search ↗
Query-Aware Knowledge Retrieval via Hyperbolic Structuring
ACL 2026
·
Chuang Zhou
DBLP profile ↗
ORCID search ↗
DREAM: Document Reconstruction via End-to-end Autoregressive Model
ACM MM 2025
·
Xin Li
DBLP profile ↗
ORCID search ↗
DS-VLM: Diffusion Supervision Vision Language Model
ICML 2025
·
Zhen Sun
DBLP profile ↗
ORCID search ↗
FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
ICML 2025
·
Zhen Sun
DBLP profile ↗
ORCID search ↗
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
ICML 2025
·
Xiong Wang
DBLP profile ↗
ORCID search ↗
Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
NeurIPS 2025
·
Yulei Qin
DBLP profile ↗
ORCID search ↗
LTD-Bench: Evaluating Large Language Models by Letting Them Draw
NeurIPS 2025
·
Liuhao Lin
DBLP profile ↗
ORCID search ↗
Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
ICLR 2025
·
Chenyu Zhou
DBLP profile ↗
ORCID search ↗
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
NeurIPS 2025
·
Chaoyou Fu
DBLP profile ↗
ORCID search ↗
Probability-Density-aware Semi-supervised Learning
AAAI 2025
·
Shuyang Liu
DBLP profile ↗
ORCID search ↗
RocketEval: Efficient automated LLM evaluation via grading checklist
ICLR 2025
·
Tianjun Wei
DBLP profile ↗
ORCID search ↗
RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing and Instruction-Following
ACL 2025
·
Junru Lu
DBLP profile ↗
ORCID search ↗
RolePlot: A Systematic Framework for Evaluating and Enhancing the Plot-Progression Capabilities of Role-Playing Agents
ACL 2025
·
Pinyi Zhang
DBLP profile ↗
ORCID search ↗
Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long Contexts
EMNLP 2025
·
Yifei Yu
DBLP profile ↗
ORCID search ↗
Tell Me What You Don't Know: Enhancing Refusal Capabilities of Role-Playing Agents via Representation Space Analysis and Editing
ACL 2025
·
Wenhao Liu
DBLP profile ↗
ORCID search ↗
Towards Universal Perception through Language-Guided Open-World Object Detection
ACM MM 2025
·
Zihan Wang
DBLP profile ↗
ORCID search ↗
TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and Speedup
NeurIPS 2025
·
Fanxu Meng
DBLP profile ↗
ORCID search ↗
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
NeurIPS 2025
·
Chaoyou Fu
DBLP profile ↗
ORCID search ↗
VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model
NeurIPS 2025
·
Zuwei Long
DBLP profile ↗
ORCID search ↗
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
CVPR 2025
·
Chaoyou Fu
DBLP profile ↗
ORCID search ↗
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
NeurIPS 2025
·
Xudong Li
DBLP profile ↗
ORCID search ↗
A General and Efficient Training for Transformer via Token Expansion
CVPR 2024
·
Wenxuan Huang
DBLP profile ↗
ORCID search ↗
Aligning and Prompting Everything All at Once for Universal Visual Perception
CVPR 2024
·
Yunhang Shen
DBLP profile ↗
ORCID search ↗
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
ACM MM 2024
·
Timin Gao
DBLP profile ↗
ORCID search ↗
Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence
EMNLP 2024
·
Junru Lu
DBLP profile ↗
ORCID search ↗
Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language Models
CVPR 2024
·
Xin Li
DBLP profile ↗
ORCID search ↗
Grab What You Need: Rethinking Complex Table Structure Recognition with Flexible Components Deliberation
AAAI 2024
·
Hao Liu
DBLP profile ↗
ORCID search ↗
HRVDA: High-Resolution Visual Document Assistant
CVPR 2024
·
Chaohu Liu
DBLP profile ↗
ORCID search ↗
Multimodal Inplace Prompt Tuning for Open-set Object Detection
ACM MM 2024
·
Guilin Li
DBLP profile ↗
ORCID search ↗
Multimodal Label Relevance Ranking via Reinforcement Learning
ECCV 2024
·
Taian Guo
DBLP profile ↗
ORCID search ↗
SPD-DDPM: Denoising Diffusion Probabilistic Models in the Symmetric Positive Definite Space
AAAI 2024
·
Yunchen Li
DBLP profile ↗
ORCID search ↗
SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger
AAAI 2024
·
Yuting Gao
DBLP profile ↗
ORCID search ↗
Visual Hallucination Elevates Speech Recognition
AAAI 2024
·
Fang Zhang
DBLP profile ↗
ORCID search ↗