PPaperPicks

Peng Gao

Shanghai Artificial Intelligence Laboratory, OpenGVLab, Shanghai, China

48 papers at tracked venues · 36 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
  2. Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
  3. CAML: Collaborative Auxiliary Modality Learning for Multi-Agent Systems
  4. Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
  5. EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
  6. EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
  7. Fontanimate: High Quality Few-Shot Font Generation Via Animating Font Transfer Process
  8. From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
  9. Let's Verify and Reinforce Image Generation Step by Step
  10. LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding
  11. Lumina-Image 2.0: a Unified and Efficient Image Generative Framework
  12. Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation
    ICLR 2025 · Peng Gao
  13. MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
  14. MMCD: Multi-Modal Collaborative Decision-Making for Connected Autonomy with Knowledge Distillation
  15. MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
  16. MMSearch: Unveiling the Potential of Large Models as Multi-modal Search Engines
  17. PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
  18. SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
  19. Spatial Preference Rewarding for MLLMs Spatial Understanding
  20. TAR3D: Creating High-Quality 3D Assets Via Next-Part Prediction
  21. UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
  22. VisualCloze: A Universal Image Generation Framework via Visual in-Context Learning
  23. Any2Point: Empowering Any-Modality Large Models for Efficient 3D Understanding
  24. BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
  25. Bridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation Skill
  26. ChartAssistant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
  27. Collaborative Decision-Making Using Spatiotemporal Graphs in Connected Autonomy
    ICRA 2024 · Peng Gao
  28. Digital Life Project: Autonomous 3D Characters with Social Intelligence
  29. FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
  30. InstructSpeech: Following Speech Editing Instructions via Large Language Models
  31. LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero-initialized Attention
  32. Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiT
  33. MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
  34. MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
  35. MTG: Mapless Trajectory Generator with Traversability Coverage for Outdoor Navigation
  36. ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
  37. Masked AutoDecoder is Effective Multi-Task Vision Generalist
  38. No Time to Train: Empowering Non-Parametric Networks for Few-Shot 3D Scene Segmentation
  39. OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
  40. OneLLM: One Framework to Align All Modalities with Language
  41. Personalize Segment Anything Model with One Shot
  42. Phased Consistency Models
  43. Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation
  44. SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
  45. SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models
  46. SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
  47. SpatialFormer: Towards Generalizable Vision Transformers with Explicit Spatial Understanding
  48. Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models