P
PaperPicks
Conferences
Renrui Zhang
54 papers at tracked venues · 48 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
NeurIPS
×11
CVPR
×10
ICLR
×9
AAAI
×7
ICCV
×5
ICML
×4
ECCV
×3
ACL
×1
ACM MM
×1
EMNLP
×1
ICRA
×1
WACV
×1
Frequent coauthors
Dongzhi Jiang
DBLP profile ↗
ORCID search ↗
×4
Jiaming Liu
DBLP profile ↗
ORCID search ↗
×3
Weifeng Lin
DBLP profile ↗
ORCID search ↗
×2
Shilin Yan
DBLP profile ↗
ORCID search ↗
×2
Zihao Deng
DBLP profile ↗
ORCID search ↗
×1
Zilu Guo
DBLP profile ↗
ORCID search ↗
×1
Victor Shea-Jay Huang
DBLP profile ↗
ORCID search ↗
×1
Sixiang Chen
DBLP profile ↗
ORCID search ↗
×1
Pengxiang Li
DBLP profile ↗
ORCID search ↗
×1
Qizhe Zhang
DBLP profile ↗
ORCID search ↗
×1
Tianshuo Peng
DBLP profile ↗
ORCID search ↗
×1
Guankun Wang
DBLP profile ↗
ORCID search ↗
×1
Papers
NL2CA: Auto-formalizing Cognitive Decision-Making from Natural Language Using an Unsupervised CriticNL2LTL Framework
AAAI 2026
·
Zihao Deng
DBLP profile ↗
ORCID search ↗
PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models
WACV 2026
·
Zilu Guo
DBLP profile ↗
ORCID search ↗
TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
AAAI 2026
·
Victor Shea-Jay Huang
DBLP profile ↗
ORCID search ↗
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
NeurIPS 2025
·
Sixiang Chen
DBLP profile ↗
ORCID search ↗
Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking
NeurIPS 2025
·
Pengxiang Li
DBLP profile ↗
ORCID search ↗
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
ICCV 2025
·
Qizhe Zhang
DBLP profile ↗
ORCID search ↗
Chimera: Improving Generalist Model with Domain-Specific Experts
ICCV 2025
·
Tianshuo Peng
DBLP profile ↗
ORCID search ↗
CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
ACM MM 2025
·
Guankun Wang
DBLP profile ↗
ORCID search ↗
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
NeurIPS 2025
·
Chengzhuo Tong
DBLP profile ↗
ORCID search ↗
Detect Anything 3D in the Wild
ICCV 2025
·
Hanxue Zhang
DBLP profile ↗
ORCID search ↗
Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning
NeurIPS 2025
·
Hao Chen
DBLP profile ↗
ORCID search ↗
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
ICCV 2025
·
Le Zhuo
DBLP profile ↗
ORCID search ↗
LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
ICLR 2025
·
Feng Li
DBLP profile ↗
ORCID search ↗
Let's Verify and Reinforce Image Generation Step by Step
CVPR 2025
·
Renrui Zhang
LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding
AAAI 2025
·
Senqiao Yang
DBLP profile ↗
ORCID search ↗
Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation
CVPR 2025
·
Yueru Jia
DBLP profile ↗
ORCID search ↗
Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation
ICLR 2025
·
Peng Gao
DBLP profile ↗
ORCID search ↗
MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
ICLR 2025
·
Renrui Zhang
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
NeurIPS 2025
·
Xinyan Chen
DBLP profile ↗
ORCID search ↗
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
AAAI 2025
·
Jiaze Wang
DBLP profile ↗
ORCID search ↗
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
ICML 2025
·
Dongzhi Jiang
DBLP profile ↗
ORCID search ↗
MMSearch: Unveiling the Potential of Large Models as Multi-modal Search Engines
ICLR 2025
·
Dongzhi Jiang
DBLP profile ↗
ORCID search ↗
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
NeurIPS 2025
·
Weifeng Lin
DBLP profile ↗
ORCID search ↗
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
ICLR 2025
·
Weifeng Lin
DBLP profile ↗
ORCID search ↗
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
ACL 2025
·
Ziyu Guo
DBLP profile ↗
ORCID search ↗
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
NeurIPS 2025
·
Dongzhi Jiang
DBLP profile ↗
ORCID search ↗
TAR3D: Creating High-Quality 3D Assets Via Next-Part Prediction
ICCV 2025
·
Xuying Zhang
DBLP profile ↗
ORCID search ↗
UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
NeurIPS 2025
·
Ruichuan An
DBLP profile ↗
ORCID search ↗
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
CVPR 2025
·
Chaoyou Fu
DBLP profile ↗
ORCID search ↗
What We Miss Matters: Learning from the Overlooked in Point Cloud Transformers
NeurIPS 2025
·
Yi Wang
DBLP profile ↗
ORCID search ↗
Cloud-Device Collaborative Learning for Multimodal Large Language Models
CVPR 2024
·
Guanqun Wang
DBLP profile ↗
ORCID search ↗
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
NeurIPS 2024
·
Dongzhi Jiang
DBLP profile ↗
ORCID search ↗
Continual-MAE: Adaptive Distribution Masked Autoencoders for Continual Test-Time Adaptation
CVPR 2024
·
Jiaming Liu
DBLP profile ↗
ORCID search ↗
FM-OV3D: Foundation Model-Based Cross-Modal Knowledge Blending for Open-Vocabulary 3D Detection
AAAI 2024
·
Dongmei Zhang
DBLP profile ↗
ORCID search ↗
Gradient-based Parameter Selection for Efficient Fine-Tuning
CVPR 2024
·
Zhi Zhang
DBLP profile ↗
ORCID search ↗
LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero-initialized Attention
ICLR 2024
·
Renrui Zhang
MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
ECCV 2024
·
Renrui Zhang
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
ICML 2024
·
Kaining Ying
DBLP profile ↗
ORCID search ↗
ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation
CVPR 2024
·
Xiaoqi Li
DBLP profile ↗
ORCID search ↗
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
ICLR 2024
·
Ke Wang
DBLP profile ↗
ORCID search ↗
NTO3D: Neural Target Object 3D Reconstruction with Segment Anything
CVPR 2024
·
Xiaobao Wei
DBLP profile ↗
ORCID search ↗
No Time to Train: Empowering Non-Parametric Networks for Few-Shot 3D Scene Segmentation
CVPR 2024
·
Xiangyang Zhu
DBLP profile ↗
ORCID search ↗
OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning
CVPR 2024
·
Lingyi Hong
DBLP profile ↗
ORCID search ↗
PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation
ECCV 2024
·
Shilin Yan
DBLP profile ↗
ORCID search ↗
Parsing All Adverse Scenes: Severity-Aware Semantic Segmentation with Mask-Enhanced Cross-Domain Consistency
AAAI 2024
·
Fuhao Li
DBLP profile ↗
ORCID search ↗
Personalize Segment Anything Model with One Shot
ICLR 2024
·
Renrui Zhang
Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation
AAAI 2024
·
Shilin Yan
DBLP profile ↗
ORCID search ↗
RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
ICRA 2024
·
Mingjie Pan
DBLP profile ↗
ORCID search ↗
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
NeurIPS 2024
·
Jiaming Liu
DBLP profile ↗
ORCID search ↗
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
ICML 2024
·
Dongyang Liu
DBLP profile ↗
ORCID search ↗
SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models
ECCV 2024
·
Ziyi Lin
DBLP profile ↗
ORCID search ↗
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
ICML 2024
·
Xudong Lu
DBLP profile ↗
ORCID search ↗
Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models
EMNLP 2024
·
Shitian Zhao
DBLP profile ↗
ORCID search ↗
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
ICLR 2024
·
Jiaming Liu
DBLP profile ↗
ORCID search ↗