PPaperPicks

Gen Luo

19 papers at tracked venues · 18 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Earth-Adapter: Bridge the Geospatial Domain Gaps with a Frequency-Guided Mixture of Adapters
  2. DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehension
  3. Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
    ICLR 2025 · Gen Luo
  4. FlashSloth : Lightning Multimodal Large Language Models via Embedded Visual Compression
  5. Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
    CVPR 2025 · Gen Luo
  6. NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
  7. Spotlight Attention: Towards Efficient LLM Generation via Non-linear Hashing-based KV Cache Retrieval
  8. Training Long-Context LLMs Efficiently via Chunk-wise Optimization
  9. WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation
  10. γ-MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
  11. 3D-GRES: Generalized 3D Referring Expression Segmentation
  12. 3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression Segmentation
  13. APL: Anchor-Based Prompt Learning for One-Stage Weakly Supervised Referring Expression Comprehension
  14. CaM: Cache Merging for Memory-efficient LLMs Inference
  15. ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
  16. Deep Instruction Tuning for Segment Anything Model
  17. Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
  18. QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual Grounding
  19. RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation