PPaperPicks

Yali Wang

Chinese Academy of Sciences, Shenzhen Institute of Advanced Technology, Guangdong-Hong Kong-Macao Joint Laboratory of Human-Machine Intelligence-Synergy Systems, China

27 papers at tracked venues · 25 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. G-UBS: Towards Robust Understanding of Implicit Feedback via Group-Aware User Behavior Simulation
  2. VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
  3. VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
  4. When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
  5. Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
  6. CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
  7. H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving
  8. LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
  9. Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
  10. Muses: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
  11. Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
  12. TimeStep Master: Asymmetrical Mixture of Timestep LoRA Experts for Versatile and Efficient Diffusion Models in Vision
  13. TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
  14. V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
  15. VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
  16. VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
  17. WeGen: A Unified Model for Interactive Multimodal Generation as We Chat
  18. EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
  19. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
  20. InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
  21. M-BEV: Masked BEV Perception for Robust Autonomous Driving
  22. MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
  23. MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
  24. SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction
  25. TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration
  26. VideoMamba: State Space Model for Efficient Video Understanding
  27. Vlogger: Make Your Dream A Vlog