PPaperPicks

Ming Yang

Ant Group, Seattle, WA, USA

27 papers at tracked venues · 23 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. SCAN: Self-Calibrated AutoregressioN for High-Quality Visual Generation
  2. Animate-X: Universal Character Image Animation with Enhanced Motion Representation
  3. CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for Guidance
  4. Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
  5. DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
  6. HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography Estimation
  7. Mimir: Improving Video Diffusion Models for Precise Text Understanding
  8. MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
  9. Reversing Flow for Image Restoration
  10. SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
  11. SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling
  12. Unified Visual Generation via Next-Set Prediction in Continuous Domain
  13. VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions
  14. Versatile Multimodal Controls for Expressive Talking Human Animation
  15. Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
  16. EVE: Efficient Zero-Shot Text-Based Video Editing With Depth Map Guidance and Temporal Consistency Constraints
  17. EcoMatcher: Efficient Clustering Oriented Matcher for Detector-Free Image Matching
  18. Learning Dynamic Tetrahedra for High-Quality Talking Head Synthesis
  19. M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval
  20. POA: Pre-training Once for Models of All Sizes
  21. Parameter-Efficient Complementary Expert Learning for Long-Tailed Visual Recognition
  22. Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
  23. Referencing Where to Focus: Improving Visual Grounding with Referential Query
  24. SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery
  25. StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models
  26. SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
  27. Towards Better Vision-Inspired Vision-Language Models