PPaperPicks

Yunchao Wei

40 papers at tracked venues · 35 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models
  2. A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
  3. Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning
  4. C2P-CLIP: Injecting Category Common Prompt in CLIP to Enhance Generalization in Deepfake Detection
  5. CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting
  6. CharaConsist: Fine-Grained Consistent Character Generation
  7. ClassDiffusion: More Aligned Personalization Tuning with Explicit Class Guidance
  8. CoMBO: Conflict Mitigation via Branched Optimization for Class Incremental Segmentation
  9. Collapsed Language Models Promote Fairness
  10. DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing
  11. Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans?
  12. FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
  13. Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention
  14. Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
  15. Memory Efficient Matting with Adaptive Token Routing
  16. NTClick: Achieving Precise Interactive Segmentation With Noise-tolerant Clicks
  17. PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild
  18. PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
  19. ReCot: Reflective Self-Correction Training for Mitigating Confirmation Bias in Large Vision-Language Models
  20. TiP4GEN: Text to Immersive Panorama 4D Scene Generation
  21. VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
  22. Visual Relation Diffusion for Human-Object Interaction Detection
  23. Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
  24. Bridge the Points: Graph-based Few-shot Segment Anything Semantically
  25. Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
  26. Diffusion for Natural Image Matting
  27. Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion Models
  28. DreamLCM: Towards High Quality Text-to-3D Generation via Latent Consistency Model
  29. Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data
  30. Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection
  31. Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning
  32. Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic Segmentation
  33. One-shot In-context Part Segmentation
  34. PVUW 2024 Challenge on Complex Video Understanding: Methods and Results
  35. PixelLM: Pixel Reasoning with Large Multimodal Model
  36. Region-Adaptive Transform with Segmentation Prior for Image Compression
  37. Region-Native Visual Tokenization
  38. Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection
  39. Segment Anything with Precise Interaction
  40. Transferable and Principled Efficiency for Open-Vocabulary Segmentation