PPaperPicks

Shaohui Lin

24 papers at tracked venues · 22 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
  2. AugKD: Ingenious Augmentations Empower Knowledge Distillation for Image Super-Resolution
  3. Complete Chess Games Enable LLM Become A Chess Master
  4. Dynamic Contrastive Knowledge Distillation for Efficient Image Restoration
  5. Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
  6. IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval
  7. Knowledge Distillation with Multi-granularity Mixture of Priors for Image Super-Resolution
  8. Probability-Density-aware Semi-supervised Learning
  9. SET: Spectral Enhancement for Tiny Object Detection
  10. TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
  11. Towards Universal Perception through Language-Guided Open-World Object Detection
  12. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
  13. WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
  14. Weakly Supervised Semantic Segmentation via Progressive Confidence Region Expansion
  15. A General and Efficient Training for Transformer via Token Expansion
  16. AQ-DETR: Low-Bit Quantized Detection Transformer with Auxiliary Queries
  17. Aligning and Prompting Everything All at Once for Universal Visual Perception
  18. CLIP in Mirror: Disentangling text from visual images through reflection
  19. CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive Learning
  20. Kumaraswamy Wavelet for Heterophilic Scene Graph Generation
  21. Rethinking Centered Kernel Alignment in Knowledge Distillation
  22. SPD-DDPM: Denoising Diffusion Probabilistic Models in the Symmetric Positive Definite Space
  23. The Ninth NTIRE 2024 Efficient Super-Resolution Challenge Report
  24. Weakly Supervised Open-Vocabulary Object Detection