PPaperPicks

Ruimao Zhang

Sun Yat-sen University, Shenzhen Campus, China

23 papers at tracked venues · 15 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Chain-of-Imagination for Reliable Instruction Following in Decision Making
  2. DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation
  3. GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion Policies
  4. High-Dynamic Radar Sequence Prediction for Weather Nowcasting Using Spatiotemporal Coherent Gaussian Representation
  5. NavigateDiff: Visual Predictors are Zero-Shot Navigation Assistants
  6. RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
  7. ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model
  8. Semantic-Supervised Spatial-Temporal Fusion for LiDAR-Based 3D Object Detection
  9. Unlock the Power of Unlabeled Data in Language Driving Model
  10. WorldSimBench: Towards Video Generation Models as World Simulators
  11. Advancing Medical Radiograph Representation Learning: A Hybrid Pre-training Paradigm with Multilevel Semantic Granularity
  12. Enhancing Human-AI Collaboration Through Logic-Guided Reasoning
  13. F-HOI: Toward Fine-Grained Semantic-Aligned 3D Human-Object Interactions
  14. FreeMan: Towards Benchmarking 3D Human Pose Estimation Under Real-World Conditions
  15. HumanTOMATO: Text-aligned Whole-body Motion Generation
  16. KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension
  17. MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
  18. Open-World Human-Object Interaction Detection via Multi-Modal Prompts
  19. SEED-Bench: Benchmarking Multimodal Large Language Models
  20. SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models
  21. Toward Accurate Camera-based 3D Object Detection via Cascade Depth Estimation and Calibration
  22. X-Pose: Detecting Any Keypoints
  23. X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge Transfer