PPaperPicks

Wentao Liu

SenseTime Research

16 papers at tracked venues · 12 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision Tasks
  2. F-LMM: Grounding Frozen Large Multimodal Models
  3. Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
  4. NADER: Neural Architecture Design via Multi-Agent Collaboration
  5. ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
  6. Unsupervised Continual Domain Shift Learning with Multi-Prototype Modeling
  7. CLIM: Contrastive Language-Image Mosaic for Region Representation
  8. CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
  9. GKGNet: Group K-Nearest Neighbor Based Graph Convolutional Network for Multi-label Image Recognition
  10. KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension
  11. Leveraging Frame Affinity for sRGB-to-RAWVideo De-Rendering
  12. PROGRAM: PROtotype GRAph Model based Pseudo-Label Learning for Test-Time Adaptation
  13. Prior Metadata-Driven RAW Reconstruction: Eliminating the Need for Per-Image Metadata
  14. UniFS: Universal Few-Shot Instance Perception with Point Representations
  15. When Pedestrian Detection Meets Multi-modal Learning: Generalist Model and Benchmark Dataset
  16. You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-person Multi-task Human-Centric Perception