PPaperPicks

Si Liu

Beihang University, School of Computer Science and Engineering, Beijing Key Laboratory of Digital Media, China

47 papers at tracked venues · 38 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AerialVLA: A Vision-Language-Action Model for Aerial Navigation with Online Dialogue
  2. MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
  3. VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
  4. 'Hi AirStar, Guide Me to the Badminton Court.'
  5. AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
  6. CoST: Efficient Collaborative Perception from Unified Spatiotemporal Perspective
  7. DOMR: Establishing Cross-View Segmentation via Dense Object Matching
  8. FlexDrive: Toward Trajectory Flexibility in Driving Scene Gaussian Splatting Reconstruction and Rendering
  9. GaussianPainter: Painting Point Cloud into 3D Gaussians with Normal Guidance
  10. Generative Map Priors for Collaborative BEV Semantic Segmentation
  11. Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
  12. LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge Distillation
  13. LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
  14. Mixture Compressor for Mixture-of-Experts LLMs Gains More
  15. Point Cluster: A Compact Message Unit for Communication-Efficient Collaborative Perception
  16. RATopo: Improving Lane Topology Reasoning via Redundancy Assignment
  17. Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
  18. RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
  19. RoboSoft'25: The 1st International Workshop on Vision-Language in Soft Robot
  20. Towards Realistic Earth-Observation Constellation Scheduling: Benchmark and Methodology
  21. Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
  22. UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning
  23. Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
  24. ViPE: Visual Perception in Parameter Space for Efficient Video-Language Understanding
  25. Video2BEV: Transforming Drone Videos to BEVs for Video-Based Geo-Localization
  26. VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
  27. Asynchronous Large Language Model Enhanced Planner for Autonomous Driving
  28. AutoVP: An Automated Visual Prompting Framework and Benchmark
  29. Collaborative Training of Tiny-Large Vision Language Models
  30. Controllable Navigation Instruction Generation with Chain of Thought Prompting
  31. CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics
  32. Customize your NeRF: Adaptive Source Driven 3D Scene Editing via Local-Global Iterative Training
  33. Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
  34. EASE-DETR: Easing the Competition among Object Queries
  35. Eliminating Cross-modal Conflicts in BEV Space for LiDAR-Camera 3D Object Detection
  36. FouriScale: A Frequency Perspective on Training-Free High-Resolution Image Synthesis
  37. GPD-VVTO: Preserving Garment Details in Video Virtual Try-On
  38. Global-Local Collaborative Inference with LLM for Lidar-Based Open-Vocabulary Detection
  39. Image Understanding Makes for A Good Tokenizer for Image Generation
  40. LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction
  41. Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection
  42. Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiT
  43. Mask-Enhanced Segment Anything Model for Tumor Lesion Semantic Segmentation
  44. Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
  45. ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target Simulation
  46. Realistic Rainy Weather Simulation for LiDARs in CARLA Simulator
  47. SAFDNet: A Simple and Effective Network for Fully Sparse 3D Object Detection