PPaperPicks

Tingting Gao

24 papers at tracked venues · 23 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Beyond Tokens: Dynamic Latent Reasoning via Semantic Residual Refinement
  2. Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
  3. IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation
  4. Live-Aid: A Large-Scale Dialogue Dataset and Benchmark for Interleaved Multi-party Interactions in Live Streaming
  5. TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs
  6. Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
  7. Why Can Distillation Work with Limited Resources? A Systematic Study
  8. CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
  9. Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models
  10. Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
  11. GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
  12. Libra-Merging: Importance-redundancy and Pruning-merging Trade-off for Acceleration Plug-in in Large Vision-Language Model
  13. LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
  14. MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
  15. MUSE: Multi-Subject Unified Synthesis Via Explicit Layout Semantic Expansion
  16. SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
  17. Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model
  18. StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
  19. TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
  20. VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform
  21. iMOVE : Instance-Motion-Aware Video Understanding
  22. Decouple Content and Motion for Conditional Image-to-Video Generation
  23. DragAnything: Motion Control for Anything Using Entity Representation
  24. Learning Multi-Dimensional Human Preference for Text-to-Image Generation