P
PaperPicks
Conferences
Kevin Qinghong Lin
16 papers at tracked venues · 15 at CORE A* · active 2024–2025
DBLP profile ↗
ORCID search ↗
Venues
CVPR
×6
NeurIPS
×4
ACM MM
×2
AAAI
×1
ECCV
×1
ICLR
×1
ICML
×1
Frequent coauthors
Qinchen Wu
DBLP profile ↗
ORCID search ↗
×1
Weijia Wu
DBLP profile ↗
ORCID search ↗
×1
Wei Pang
DBLP profile ↗
ORCID search ↗
×1
Yuchao Gu
DBLP profile ↗
ORCID search ↗
×1
Jinheng Xie
DBLP profile ↗
ORCID search ↗
×1
Jiaqi Wang
DBLP profile ↗
ORCID search ↗
×1
Shravan Nayak
DBLP profile ↗
ORCID search ↗
×1
Muhammet Furkan Ilaslan
DBLP profile ↗
ORCID search ↗
×1
Difei Gao
DBLP profile ↗
ORCID search ↗
×1
Ziteng Gao
DBLP profile ↗
ORCID search ↗
×1
Shiwei Wu
DBLP profile ↗
ORCID search ↗
×1
Joya Chen
DBLP profile ↗
ORCID search ↗
×1
Papers
GUI-Narrator: Detecting and Captioning Computer GUI Actions
ACM MM 2025
·
Qinchen Wu
DBLP profile ↗
ORCID search ↗
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
CVPR 2025
·
Weijia Wu
DBLP profile ↗
ORCID search ↗
Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
NeurIPS 2025
·
Wei Pang
DBLP profile ↗
ORCID search ↗
ROICtrl: Boosting Instance Control for Visual Generation
CVPR 2025
·
Yuchao Gu
DBLP profile ↗
ORCID search ↗
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
ICLR 2025
·
Jinheng Xie
DBLP profile ↗
ORCID search ↗
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
CVPR 2025
·
Kevin Qinghong Lin
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
NeurIPS 2025
·
Jiaqi Wang
DBLP profile ↗
ORCID search ↗
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
ICML 2025
·
Shravan Nayak
DBLP profile ↗
ORCID search ↗
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
AAAI 2025
·
Muhammet Furkan Ilaslan
DBLP profile ↗
ORCID search ↗
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
CVPR 2025
·
Kevin Qinghong Lin
AssistEditor: Multi-Agent Collaboration for GUI Workflow Automation in Video Creation
ACM MM 2024
·
Difei Gao
DBLP profile ↗
ORCID search ↗
Bootstrapping SparseFormers from Vision Foundation Models
CVPR 2024
·
Ziteng Gao
DBLP profile ↗
ORCID search ↗
Learning Video Context as Interleaved Multimodal Sequences
ECCV 2024
·
Kevin Qinghong Lin
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
NeurIPS 2024
·
Kevin Qinghong Lin
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
NeurIPS 2024
·
Shiwei Wu
DBLP profile ↗
ORCID search ↗
VideoLLM-online: Online Video Large Language Model for Streaming Video
CVPR 2024
·
Joya Chen
DBLP profile ↗
ORCID search ↗