PPaperPicks

Yun Zheng

11 papers at tracked venues · 9 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Aligned Better, Listen Better for Audio-Visual Large Language Models
  2. CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
  3. ContextHOI: Spatial Context Learning for Human-Object Interaction Detection
  4. DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
  5. Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
  6. Orchestrating the Symphony of Prompt Distribution Learning for Human-Object Interaction Detection
  7. UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface
  8. CoReS: Orchestrating the Dance of Reasoning and Segmentation
  9. CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-Training
  10. FuseTeacher: Modality-Fused Encoders are Strong Vision Supervisors
  11. Relevant Intrinsic Feature Enhancement Network for Few-Shot Semantic Segmentation