PPaperPicks

Lei Zhang

International Digital Economy Academy (IDEA), Shenzhen, China

28 papers at tracked venues · 22 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
  2. SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
  3. T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection
  4. Towards Better Code Understanding in Decoder-Only Models with Contrastive Learning
  5. Adversarial Diffusion Compression for Real-World Image Super-Resolution
  6. Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive Injection
  7. CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
  8. HandOS: 3D Hand Reconstruction in One Stage
  9. HumanMM: Global Human Motion Recovery from Multi-shot Videos
  10. LeanGaussian: Breaking Pixel or Point Cloud Correspondence in Modeling 3D Gaussians
  11. OSMamba: Omnidirectional Spectral Mamba with Dual-Domain Prior Generator for Exposure Correction
  12. Open-Set Image Tagging with Multi-Grained Text Supervision
  13. Scaling Speech-Text Pre-training with Synthetic Interleaved Data
  14. SkillMimic: Learning Basketball Interaction Skills from Demonstrations
  15. UniGS: Modeling Unitary 3D Gaussians for Novel View Synthesis from Sparse-View Images
  16. Compress3D: A Compressed Latent Space for 3D Generation from a Single Image
  17. DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation
  18. Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection
  19. HumanTOMATO: Text-aligned Whole-body Motion Generation
  20. LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
  21. Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic Prompts
  22. Recognize Anything: A Strong Image Tagging Model
  23. Segment and Recognize Anything at Any Granularity
  24. T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
  25. TAPTR: Tracking Any Point with Transformers as Detection
  26. TOSS: High-quality Text-guided Novel View Synthesis from a Single Image
  27. Tag2Text: Guiding Vision-Language Model via Image Tagging
  28. Visual in-Context Prompting