PPaperPicks

Lewei Lu

26 papers at tracked venues · 24 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Docopilot: Improving Multimodal Models for Document-Level Understanding
  2. Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
  3. GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
  4. HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
  5. MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
  6. NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
  7. PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
  8. Spatial Preference Rewarding for MLLMs Spatial Understanding
  9. Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM
  10. SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
  11. Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
  12. ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
  13. Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
  14. ControlLLM: Augment Language Models with Tools by Searching on Graphs
  15. Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
  16. Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
  17. LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors
  18. Learning 1D Causal Visual Representation with De-focus Attention Networks
  19. Masked AutoDecoder is Effective Multi-Task Vision Generalist
  20. Modeling Continuous Motion for 3D Point Cloud Object Tracking
  21. Needle In A Multimodal Haystack
  22. Parameter-Inverted Image Pyramid Networks
  23. The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
  24. Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
  25. VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
  26. Weakly Supervised Monocular 3D Detection with a Single-View Image