PPaperPicks

Jingdong Chen

28 papers at tracked venues · 23 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMs
  2. SCAN: Self-Calibrated AutoregressioN for High-Quality Visual Generation
  3. UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
  4. ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
  5. Animate-X: Universal Character Image Animation with Enhanced Motion Representation
  6. CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for Guidance
  7. Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
  8. HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography Estimation
  9. Mimir: Improving Video Diffusion Models for Precise Text Understanding
  10. MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
  11. On the Design of a Robust Superdirective Beamformer and Topology Parameter Optimization with Frustum-Shaped Microphone Arrays Featuring Multiple Rings
  12. Reversing Flow for Image Restoration
  13. SkySense V2: A Unified Foundation Model for Multi-Modal Remote Sensing
  14. SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling
  15. The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
  16. VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations
  17. VideoMAR: Autoregressive Video Generation with Continuous Tokens
  18. When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
  19. Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
  20. EcoMatcher: Efficient Clustering Oriented Matcher for Detector-Free Image Matching
  21. Large Multimodal Model Compression via Iterative Efficient Pruning and Distillation
  22. Learning Dynamic Tetrahedra for High-Quality Talking Head Synthesis
  23. LogicMP: A Neuro-symbolic Approach for Encoding First-order Logic Constraints
  24. POA: Pre-training Once for Models of All Sizes
  25. Parameter-Efficient Complementary Expert Learning for Long-Tailed Visual Recognition
  26. SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery
  27. StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models
  28. Towards Better Vision-Inspired Vision-Language Models