PPaperPicks

Xin Tan

East China Normal University, Shanghai, China

31 papers at tracked venues · 28 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Diffusion Implicit Policy for Unpaired Scene-aware Motion Synthesis
  2. Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
  3. LidarPainter: One-Step Away from Any Lidar View to Novel Guidance
  4. Multi-Step Deformable Gaussian Splatting for Dynamic Scene Rendering
  5. NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
  6. Zero-Shot Robotic Manipulation via 3D Gaussian Splatting-Enhanced Multimodal Retrieval-Augmented Generation
  7. DrivingForward: Feed-forward 3D Gaussian Splatting for Driving Scene Reconstruction from Flexible Surround-view Input
  8. EyeSeg: An Uncertainty-Aware Eye Segmentation Framework for AR/VR
  9. FastLGS: Speeding Up Language Embedded Gaussians with Feature Grid Mapping
  10. From Enhancement to Understanding: Build a Generalized Bridge for Low-Light Vision via Semantically Consistent Unsupervised Fine-Tuning
  11. IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval
  12. Large Continual Instruction Assistant
  13. MOS: Modeling Object-Scene Associations in Generalized Category Discovery
  14. One-for-More: Continual Diffusion Model for Anomaly Detection
  15. PFDepth: Heterogeneous Pinhole-Fisheye Joint Depth Estimation via Distortion-aware Gaussian-Splatted Volumetric Fusion
  16. Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models
  17. Stylized-Face: A Million-Level Stylized Face Dataset for Face Recognition
  18. Beyond the Label Itself: Latent Labels Enhance Semi-supervised Point Cloud Panoptic Segmentation
  19. Building a Strong Pre-Training Baseline for Universal 3D Large-Scale Perception
  20. COTR: Compact Occupancy TRansformer for Vision-Based 3D Occupancy Prediction
  21. CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
  22. Continuous Piecewise-Affine Based Motion Model for Image Animation
  23. Domain Alignment with Large Vision-language Models for Cross-domain Remote Sensing Image Retrieval
  24. Domain-Hallucinated Updating for Multi-Domain Face Anti-spoofing
  25. Harmonizing Visual Text Comprehension and Generation
  26. Image-text Retrieval with Main Semantics Consistency
  27. LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
  28. Learning Task-Aware Language-Image Representation for Class-Incremental Object Detection
  29. Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer
  30. Prompt Gradient Projection for Continual Learning
  31. PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection