PPaperPicks

Yi Wang

Shanghai AI Laboratory, China

19 papers at tracked venues · 17 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. LLaTiSA: Towards Difficulty-Stratified Time Series Reasoning from Visual Perception to Semantics
  2. TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
  3. VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
  4. Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
  5. DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
  6. Make Your Training Flexible: Towards Deployment-Efficient Video Models
  7. Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
  8. StreamForest: Efficient Online Video Understanding with Persistent Event Memory
  9. Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
  10. TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
  11. VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
  12. ViLLa: Video Reasoning Segmentation with Large Language Model
  13. VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
  14. Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?
  15. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
    ICLR 2024 · Yi Wang
  16. InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
    ECCV 2024 · Yi Wang
  17. MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
  18. SyncVIS: Synchronized Video Instance Segmentation
  19. VideoMamba: State Space Model for Efficient Video Understanding