PPaperPicks

Tong He

Shanghai AI Lab, Shanghai

31 papers at tracked venues · 26 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. AETHER: Geometric-Aware Unified World Modeling
  2. DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
  3. Depth Any Video with Scalable Synthetic Data
  4. EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
  5. GigaGS: 3D Gaussian Based Planar Representation for Large-Scene Surface Reconstruction
  6. GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving
  7. Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation
  8. MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers
  9. ND-SDF: Learning Normal Deflection Fields for High-Fidelity Indoor Reconstruction
  10. SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
  11. Sekai: A Video Dataset towards World Exploration
  12. Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning
  13. VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
  14. Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction
  15. Agent3D-Zero: An Agent for Zero-Shot 3D Understanding
  16. Boosting Residual Networks with Group Knowledge
  17. CaMML: Context-Aware Multimodal Learner for Large Models
  18. Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything Model
  19. DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
  20. DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware Diffusion
  21. DreamComposer: Controllable 3D Object Generation via Multi-View Conditions
  22. EMR-Merging: Tuning-Free High-Performance Model Merging
  23. Frozen CLIP Transformer Is an Efficient Point Cloud Encoder
  24. GVGEN: Text-to-3D Generation with Volumetric Representation
  25. NeuRodin: A Two-stage Framework for High-Fidelity Neural Surface Reconstruction
  26. Pixel-GS: Density Control with Pixel-Aware Gradient for 3D Gaussian Splatting
  27. Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
  28. Point Transformer V3: Simpler, Faster, Stronger
  29. PredBench: Benchmarking Spatio-Temporal Prediction Across Diverse Disciplines
  30. TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation
  31. UniPAD: A Universal Pre-Training Paradigm for Autonomous Driving