PPaperPicks

Ming-Ming Cheng

47 papers at tracked venues · 45 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection
  2. SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
  3. Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection
  4. A Simple Detector with Frame Dynamics is a Strong Tracker
  5. AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction
  6. Advancing Textual Prompt Learning with Anchored Attributes
  7. Anchor Token Matching: Implicit Structure Locking for Training-Free AR Image Editing
  8. AngleRoCL: Angle-Robust Concept Learning for Physically View-Invariant Adversarial Patches
  9. DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
  10. DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data
  11. DISTA-Net: Dynamic Closely-Spaced Infrared Small Target Unmixing
  12. DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible Patches
  13. From Words to Worth: Newborn Article Impact Prediction with LLM
  14. GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
  15. InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration
  16. KAC: Kolmogorov-Arnold Classifier for Continual Learning
  17. Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning
  18. Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
  19. OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
  20. One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
  21. RSAR: Restricted State Angle Resolver and Rotated SAR Benchmark
  22. Re-Aligning Language to Visual Objects with an Agentic Workflow
  23. Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
  24. Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
  25. Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology
  26. TAR3D: Creating High-Quality 3D Assets Via Next-Part Prediction
  27. TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
  28. Towards RAW Object Detection in Diverse Conditions
  29. Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
  30. VisualCloze: A Universal Image Generation Framework via Visual in-Context Learning
  31. Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
  32. CorrMatch: Label Propagation via Correlation Matching for Semi-Supervised Semantic Segmentation
  33. CrossKD: Cross-Head Knowledge Distillation for Object Detection
  34. DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
  35. Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
  36. Fine-Grained Knowledge Selection and Restoration for Non-exemplar Class Incremental Learning
  37. Generative Multi-modal Models are Good Class-Incremental Learners
  38. Let's Start Over: Retraining with Selective Samples for Generalized Category Discovery
  39. OPUS: Occupancy Prediction Using a Sparse Set
  40. PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
  41. SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection
  42. StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
  43. Task-Adaptive Saliency Guidance for Exemplar-Free Class Incremental Learning
  44. TeMO: Towards Text-Driven 3D Stylization for Multi-Object Meshes
  45. Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
  46. Towards Stable 3D Object Detection
  47. Traffic Scene Parsing Through the TSP6K Dataset