PPaperPicks

Kevin Qinghong Lin

16 papers at tracked venues · 15 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. GUI-Narrator: Detecting and Captioning Computer GUI Actions
  2. MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
  3. Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
  4. ROICtrl: Boosting Instance Control for Visual Generation
  5. Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
  6. ShowUI: One Vision-Language-Action Model for GUI Visual Agent
    CVPR 2025 · Kevin Qinghong Lin
  7. Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
  8. UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
  9. VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
  10. VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
    CVPR 2025 · Kevin Qinghong Lin
  11. AssistEditor: Multi-Agent Collaboration for GUI Workflow Automation in Video Creation
  12. Bootstrapping SparseFormers from Vision Foundation Models
  13. Learning Video Context as Interleaved Multimodal Sequences
    ECCV 2024 · Kevin Qinghong Lin
  14. VideoGUI: A Benchmark for GUI Automation from Instructional Videos
    NeurIPS 2024 · Kevin Qinghong Lin
  15. VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
  16. VideoLLM-online: Online Video Large Language Model for Streaming Video