PPaperPicks

Huaibo Huang

20 papers at tracked venues · 18 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. T2Agent: A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree Search
  2. Breaking the Low-Rank Dilemma of Linear Attention
  3. DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
  4. InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning
  5. MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
  6. Rectifying Magnitude Neglect in Linear Attention
  7. Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
  8. Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
  9. Vision Transformer with Sparse Scan Prior
  10. DeVAn: Dense Video Annotation for Video-Language Models
  11. DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset Curation
  12. FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMs
  13. Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content Generation
  14. Heterogeneous Test-Time Training for Multi-Modal Person Re-identification
  15. INSTASTYLE: Inversion Noise of a Stylized Image is Secretly a Style Adviser
  16. Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image Restoration
  17. RMT: Retentive Networks Meet Vision Transformers
  18. Uncertainty-Aware Source-Free Adaptive Image Super-Resolution with Wavelet Augmentation Transformer
  19. Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model
  20. ZePo: Zero-Shot Portrait Stylization with Faster Sampling