PPaperPicks

Qixiang Ye

17 papers at tracked venues · 15 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Adaptive Keyframe Sampling for Long Video Understanding
  2. Building Vision Models upon Heat Conduction
  3. ChatterBox: Multimodal Referring and Grounding with Chain-of-Questions
  4. ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
  5. DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution
  6. RS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation Model
  7. Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
  8. YOLOv12: Attention-Centric Real-Time Object Detectors
  9. Artemis: Towards Referential Understanding in Complex Videos
  10. ControlCap: Controllable Region-Level Captioning
  11. Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
  12. Grounding Multimodal Large Language Models to the World
  13. Kepler codebook
  14. Ray Denoising: Depth-Aware Hard Negative Sampling for Multi-view 3D Object Detection
  15. Regressor-Segmenter Mutual Prompt Learning for Crowd Counting
  16. Spatial Transform Decoupling for Oriented Object Detection
  17. VMamba: Visual State Space Model