PPaperPicks

Xiaoyi Dong

32 papers at tracked venues · 30 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
  2. Bootstrap3D: Improving Multi-View Diffusion Model with Synthetic Data
  3. ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
  4. Conical Visual Concentration for Efficient Large Vision-Language Models
  5. Deciphering Cross-Modal Alignment in Large Vision-Language Models Via Modality Integration Rate
  6. Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
  7. HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance
  8. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
  9. Light-a-Video: Training-Free Video Relighting via Progressive Light Fusion
  10. MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
  11. MM-IFEngine: Towards Multimodal Instruction Following
  12. Maximum Entropy Reinforcement Learning with Diffusion Policy
    ICML 2025 · Xiaoyi Dong
  13. MotionClone: Training-Free Motion Cloning for Controllable Video Generation
  14. OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
  15. SAM2LONG: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
  16. SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
  17. SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation
  18. Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
  19. VideoRoPE: What Makes for Good Video Rotary Position Embedding?
  20. Visual-RFT: Visual Reinforcement Fine-Tuning
  21. X-Prompt: Generalizable Auto-Regressive Visual Learning with In-Context Prompting
  22. Are We on the Right Way for Evaluating Large Vision-Language Models?
  23. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
    NeurIPS 2024 · Xiaoyi Dong
  24. Long-CLIP: Unlocking the Long-Text Capability of CLIP
  25. MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
  26. MMLONGBENCH-DOC: Benchmarking Long-context Document Understanding with Visualizations
  27. OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
  28. ShareGPT4V: Improving Large Multi-modal Models with Better Captions
  29. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
  30. Streaming Long Video Understanding with Large Language Models
  31. VIGC: Visual Instruction Generation and Correction
  32. VLMEvalKit: An Open-Source ToolKit for Evaluating Large Multi-Modality Models