PPaperPicks

Yilun Chen

20 papers at tracked venues · 11 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
  2. A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
  3. Bench4Merge: A Comprehensive Benchmark for Merging in Realistic Dense Traffic with Micro-Interactive Vehicles
  4. CoopDETR: A Unified Cooperative Perception Framework for 3D Detection via Object Query
  5. Dual-AEB: Synergizing Rule-Based and Multimodal Large Language Models for Effective Emergency Braking
  6. GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
  7. Language-to-Space Programming for Training-Free 3D Visual Grounding
  8. LiON: Learning Point-Wise Abstaining Penalty for LiDAR Outlier DetectioN Using Diverse Synthetic Data
  9. MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance
  10. Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
  11. RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
  12. SLM-Mod: Small Language Models Surpass LLMs at Content Moderation
  13. Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving
  14. Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
  15. EMIFF: Enhanced Multi-scale Image Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object Detection
  16. MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
  17. More Than Routing: Joint GPS and Route Modeling for Refine Trajectory Representation Learning
  18. PointLLM: Empowering Large Language Models to Understand Point Clouds
  19. TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
  20. What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights