PPaperPicks

Renrui Zhang

54 papers at tracked venues · 48 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. NL2CA: Auto-formalizing Cognitive Decision-Making from Natural Language Using an Unsupervised CriticNL2LTL Framework
  2. PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models
  3. TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
  4. AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
  5. Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking
  6. Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
  7. Chimera: Improving Generalist Model with Domain-Specific Experts
  8. CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
  9. Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
  10. Detect Anything 3D in the Wild
  11. Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning
  12. From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
  13. LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
  14. Let's Verify and Reinforce Image Generation Step by Step
    CVPR 2025 · Renrui Zhang
  15. LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding
  16. Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation
  17. Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation
  18. MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
    ICLR 2025 · Renrui Zhang
  19. MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
  20. MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
  21. MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
  22. MMSearch: Unveiling the Potential of Large Models as Multi-modal Search Engines
  23. Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
  24. PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
  25. SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
  26. T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
  27. TAR3D: Creating High-Quality 3D Assets Via Next-Part Prediction
  28. UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
  29. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
  30. What We Miss Matters: Learning from the Overlooked in Point Cloud Transformers
  31. Cloud-Device Collaborative Learning for Multimodal Large Language Models
  32. CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
  33. Continual-MAE: Adaptive Distribution Masked Autoencoders for Continual Test-Time Adaptation
  34. FM-OV3D: Foundation Model-Based Cross-Modal Knowledge Blending for Open-Vocabulary 3D Detection
  35. Gradient-based Parameter Selection for Efficient Fine-Tuning
  36. LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero-initialized Attention
    ICLR 2024 · Renrui Zhang
  37. MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
    ECCV 2024 · Renrui Zhang
  38. MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
  39. ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation
  40. MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
  41. NTO3D: Neural Target Object 3D Reconstruction with Segment Anything
  42. No Time to Train: Empowering Non-Parametric Networks for Few-Shot 3D Scene Segmentation
  43. OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning
  44. PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation
  45. Parsing All Adverse Scenes: Severity-Aware Semantic Segmentation with Mask-Enhanced Cross-Domain Consistency
  46. Personalize Segment Anything Model with One Shot
    ICLR 2024 · Renrui Zhang
  47. Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation
  48. RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
  49. RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
  50. SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
  51. SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models
  52. SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
  53. Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models
  54. ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation