PPaperPicks

Xiaohan Wang

25 papers at tracked venues · 19 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning
  2. Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization
    AAAI 2026 · Xiaohan Wang
  3. Apollo: An Exploration of Video Understanding in Large Multimodal Models
  4. Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
  5. BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
  6. Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
  7. Innovative Thinking, Infinite Humor: Humor Research of Large Language Models through Structured Thought Leaps
  8. Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
  9. LAFNET: Lightweight Aerial Fire Detection Model for Onboard Edge Computing
  10. Targeted Learning for Variable Importance
    UAI 2025 · Xiaohan Wang
  11. Video Action Differencing
  12. Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
  13. A Category Agnostic Model for Visual Rearrangment
  14. An Interactive Navigation Method with Effect-oriented Affordance
    CVPR 2024 · Xiaohan Wang
  15. Continual Multimodal Knowledge Graph Construction
  16. Cross-Sentence Gloss Consistency for Continuous Sign Language Recognition
  17. DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
  18. Describing Differences in Image Sets with Natural Language
  19. EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models
  20. Editing Conceptual Knowledge for Large Language Models
    EMNLP 2024 · Xiaohan Wang
  21. Imagine Before Go: Self-Supervised Generative Map for Object Goal Navigation
  22. Interpretable3D: An Ad-Hoc Interpretable Classifier for 3D Point Clouds
  23. Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
  24. VideoAgent: Long-Form Video Understanding with Large Language Model as Agent
    ECCV 2024 · Xiaohan Wang
  25. Why are Visually-Grounded Language Models Bad at Image Classification?