P
PaperPicks
Conferences
Yi Wang
Shanghai AI Laboratory, China
19 papers at tracked venues · 17 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
NeurIPS
×5
ICCV
×4
ICLR
×3
AAAI
×2
CVPR
×2
ECCV
×2
ACL
×1
Frequent coauthors
Rongkun Zheng
DBLP profile ↗
ORCID search ↗
×3
Xiangyu Zeng
DBLP profile ↗
ORCID search ↗
×2
Ziang Yan
DBLP profile ↗
ORCID search ↗
×2
Kunchang Li
DBLP profile ↗
ORCID search ↗
×2
Yueyang Ding
DBLP profile ↗
ORCID search ↗
×1
Meng Chu
DBLP profile ↗
ORCID search ↗
×1
Zikang Wang
DBLP profile ↗
ORCID search ↗
×1
Zun Wang
DBLP profile ↗
ORCID search ↗
×1
Jiahe Zhao
DBLP profile ↗
ORCID search ↗
×1
Chenting Wang
DBLP profile ↗
ORCID search ↗
×1
Jiashuo Yu
DBLP profile ↗
ORCID search ↗
×1
Qingsong Zhao
DBLP profile ↗
ORCID search ↗
×1
Papers
LLaTiSA: Towards Difficulty-Stratified Time Series Reasoning from Visual Perception to Semantics
ACL 2026
·
Yueyang Ding
DBLP profile ↗
ORCID search ↗
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
AAAI 2026
·
Meng Chu
DBLP profile ↗
ORCID search ↗
VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
AAAI 2026
·
Zikang Wang
DBLP profile ↗
ORCID search ↗
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
ICLR 2025
·
Zun Wang
DBLP profile ↗
ORCID search ↗
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
ICCV 2025
·
Jiahe Zhao
DBLP profile ↗
ORCID search ↗
Make Your Training Flexible: Towards Deployment-Efficient Video Models
ICCV 2025
·
Chenting Wang
DBLP profile ↗
ORCID search ↗
Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
NeurIPS 2025
·
Rongkun Zheng
DBLP profile ↗
ORCID search ↗
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
NeurIPS 2025
·
Xiangyu Zeng
DBLP profile ↗
ORCID search ↗
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
CVPR 2025
·
Ziang Yan
DBLP profile ↗
ORCID search ↗
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
ICLR 2025
·
Xiangyu Zeng
DBLP profile ↗
ORCID search ↗
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
ICCV 2025
·
Jiashuo Yu
DBLP profile ↗
ORCID search ↗
ViLLa: Video Reasoning Segmentation with Large Language Model
ICCV 2025
·
Rongkun Zheng
DBLP profile ↗
ORCID search ↗
VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
NeurIPS 2025
·
Ziang Yan
DBLP profile ↗
ORCID search ↗
Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?
NeurIPS 2024
·
Qingsong Zhao
DBLP profile ↗
ORCID search ↗
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
ICLR 2024
·
Yi Wang
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
ECCV 2024
·
Yi Wang
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
CVPR 2024
·
Kunchang Li
DBLP profile ↗
ORCID search ↗
SyncVIS: Synchronized Video Instance Segmentation
NeurIPS 2024
·
Rongkun Zheng
DBLP profile ↗
ORCID search ↗
VideoMamba: State Space Model for Efficient Video Understanding
ECCV 2024
·
Kunchang Li
DBLP profile ↗
ORCID search ↗