P
PaperPicks
Conferences
Yixiao Ge
29 papers at tracked venues · 24 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
CVPR
×12
ICCV
×5
ACL
×2
ECCV
×2
ICML
×2
ICRA
×2
AAAI
×1
ICLR
×1
NAACL
×1
NeurIPS
×1
Frequent coauthors
Yi Chen
DBLP profile ↗
ORCID search ↗
×2
Xubing Ye
DBLP profile ↗
ORCID search ↗
×2
Yuying Ge
DBLP profile ↗
ORCID search ↗
×2
Yicheng Xiao
DBLP profile ↗
ORCID search ↗
×2
Chengyue Wu
DBLP profile ↗
ORCID search ↗
×2
Ruyang Liu
DBLP profile ↗
ORCID search ↗
×2
Junhao Cheng
DBLP profile ↗
ORCID search ↗
×1
Shijie Ma
DBLP profile ↗
ORCID search ↗
×1
Rui Yang
DBLP profile ↗
ORCID search ↗
×1
Shuai Yang
DBLP profile ↗
ORCID search ↗
×1
Fengyuan Shi
DBLP profile ↗
ORCID search ↗
×1
Alessandro Fornasier
DBLP profile ↗
ORCID search ↗
×1
Papers
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
ACL 2026
·
Yi Chen
DBLP profile ↗
ORCID search ↗
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
CVPR 2025
·
Xubing Ye
DBLP profile ↗
ORCID search ↗
AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
ICCV 2025
·
Junhao Cheng
DBLP profile ↗
ORCID search ↗
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
CVPR 2025
·
Yuying Ge
DBLP profile ↗
ORCID search ↗
Equivariant Filter Design for Range-Only SLAM
ICRA 2025
·
Yixiao Ge
GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers
ICCV 2025
·
Shijie Ma
DBLP profile ↗
ORCID search ↗
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
ICML 2025
·
Rui Yang
DBLP profile ↗
ORCID search ↗
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
ICML 2025
·
Yicheng Xiao
DBLP profile ↗
ORCID search ↗
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
ICCV 2025
·
Yi Chen
DBLP profile ↗
ORCID search ↗
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
NAACL 2025
·
Chengyue Wu
DBLP profile ↗
ORCID search ↗
SEED-Story: Multimodal Long Story Generation with Large Language Model
ICCV 2025
·
Shuai Yang
DBLP profile ↗
ORCID search ↗
Scalable Image Tokenization with Index Backpropagation Quantization
ICCV 2025
·
Fengyuan Shi
DBLP profile ↗
ORCID search ↗
VoCo-LLaMA: Towards Vision Compression with Large Language Models
CVPR 2025
·
Xubing Ye
DBLP profile ↗
ORCID search ↗
An Equivariant Approach to Robust State Estimation for the ArduPilot Autopilot System
ICRA 2024
·
Alessandro Fornasier
DBLP profile ↗
ORCID search ↗
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
CVPR 2024
·
Ruyang Liu
DBLP profile ↗
ORCID search ↗
Cached Transformers: Improving Transformers with Differentiable Memory Cachde
AAAI 2024
·
Zhaoyang Zhang
DBLP profile ↗
ORCID search ↗
DreamDiffusion: High-Quality EEG-to-Image Generation with Temporal Masked Signal Modeling and CLIP Alignment
ECCV 2024
·
Yunpeng Bai
DBLP profile ↗
ORCID search ↗
LLaMA Pro: Progressive LLaMA with Block Expansion
ACL 2024
·
Chengyue Wu
DBLP profile ↗
ORCID search ↗
Low-Rank Approximation for Sparse Attention in Multi-Modal LLMs
CVPR 2024
·
Lin Song
DBLP profile ↗
ORCID search ↗
Making LLaMA SEE and Draw with SEED Tokenizer
ICLR 2024
·
Yuying Ge
DBLP profile ↗
ORCID search ↗
MambaTree: Tree Topology is All You Need in State Space Model
NeurIPS 2024
·
Yicheng Xiao
DBLP profile ↗
ORCID search ↗
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
CVPR 2024
·
Yiyuan Zhang
DBLP profile ↗
ORCID search ↗
Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis
CVPR 2024
·
Yuchao Gu
DBLP profile ↗
ORCID search ↗
SEED-Bench: Benchmarking Multimodal Large Language Models
CVPR 2024
·
Bohao Li
DBLP profile ↗
ORCID search ↗
ST-LLM: Large Language Models Are Effective Temporal Learners
ECCV 2024
·
Ruyang Liu
DBLP profile ↗
ORCID search ↗
SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models
CVPR 2024
·
Yuzhou Huang
DBLP profile ↗
ORCID search ↗
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
CVPR 2024
·
Xiaohan Ding
DBLP profile ↗
ORCID search ↗
VIT-LENS: Towards Omni-modal Representations
CVPR 2024
·
Weixian Lei
DBLP profile ↗
ORCID search ↗
YOLO-World: Real-Time Open-Vocabulary Object Detection
CVPR 2024
·
Tianheng Cheng
DBLP profile ↗
ORCID search ↗