PPaperPicks

Yuxin Peng

Peking University, Beijing, China

45 papers at tracked venues · 41 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification
  2. HD²-SSC: High-Dimension High-Density Semantic Scene Completion for Autonomous Driving
  3. Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
  4. Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
  5. Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing
  6. ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
  7. DASK: Distribution Rehearsing via Adaptive Style Kernel Learning for Exemplar-Free Lifelong Person Re-Identification
  8. DKC: Differentiated Knowledge Consolidation for Cloth-Hybrid Lifelong Person Re-identification
  9. DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
  10. EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
  11. Hierarchical Event Memory for Accurate and Low-Latency Online Video Temporal Grounding
  12. Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
  13. Interact-Custom: Customized Human Object Interaction Image Generation
  14. Investigating Domain Gaps for Indoor 3D Object Detection
  15. MAI: A Multi-turn Aggregation-Iteration Model for Composed Image Retrieval
  16. Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept Calibration
  17. PosterO: Structuring Layout Trees to Enable Language Models in Generalized Content-Aware Layout Generation
  18. SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute Prompting
  19. SPHERE: Semantic-PHysical Engaged REpresentation for 3D Semantic Scene Completion
  20. STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
  21. Scan-and-Print: Patch-level Data Summarization and Augmentation for Content-aware Layout Generation in Poster Design
  22. Selective Visual Prompting in Vision Mamba
  23. TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-Enhanced Relation-Aware Knowledge Transferring
  24. UPP: Unified Point-Level Prompting for Robust Point Cloud Analysis
  25. Comprehensive Visual Grounding for Video Description
  26. Continual Vision-Language Retrieval via Dynamic Knowledge Rectification
  27. DART: Dual-Modal Adaptive Online Prompting and Knowledge Retention for Test-Time Adaptation
  28. Distribution-Aware Knowledge Prototyping for Non-Exemplar Lifelong Person Re-Identification
  29. Exploring Conditional Multi-modal Prompts for Zero-Shot HOI Detection
  30. FCS: Feature Calibration and Separation for Non-Exemplar Class Incremental Learning
  31. FashionERN: Enhance-and-Refine Network for Composed Fashion Image Retrieval
  32. FineFMPL: Fine-grained Feature Mining Prompt Learning for Few-Shot Class Incremental Learning
  33. FinePOSE: Fine-Grained Prompt-Driven 3D Human Pose Estimation via Diffusion Models
  34. FineParser: A Fine-Grained Spatio-Temporal Action Parser for Human-Centric Action Quality Assessment
  35. FineSports: A Multi-Person Hierarchical Sports Video Dataset for Fine-Grained Action Understanding
  36. Firzen: Firing Strict Cold-Start Items with Frozen Heterogeneous and Homogeneous Graphs for Recommendation
  37. InsVP: Efficient Instance Visual Prompting from Image Itself
  38. Learning Continual Compatible Representation for Re-indexing Free Lifelong Person Re-identification
  39. Mitigate Catastrophic Remembering via Continual Knowledge Purification for Noisy Lifelong Person Re-Identification
  40. Progressive Prototype Evolving for Dual-Forgetting Mitigation in Non-Exemplar Online Continual Learning
  41. RelScene: A Benchmark and baseline for Spatial Relations in text-driven 3D Scene Generation
  42. ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
  43. SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection
  44. Semantic-Aware Human Object Interaction Image Generation
  45. Training-Free Video Temporal Grounding Using Large-Scale Pre-trained Models