PPaperPicks

Pan Zhang

Shanghai Artificial Intelligence Laboratory, Shanghai, China

31 papers at tracked venues · 29 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Bootstrap3D: Improving Multi-View Diffusion Model with Synthetic Data
  2. ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
  3. Conical Visual Concentration for Efficient Large Vision-Language Models
  4. Deciphering Cross-Modal Alignment in Large Vision-Language Models Via Modality Integration Rate
  5. Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
  6. HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance
  7. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
  8. Light-a-Video: Training-Free Video Relighting via Progressive Light Fusion
  9. MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
  10. MM-IFEngine: Towards Multimodal Instruction Following
  11. MotionClone: Training-Free Motion Cloning for Controllable Video Generation
  12. OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
  13. SAM2LONG: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
  14. SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
  15. SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation
  16. Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
  17. VideoRoPE: What Makes for Good Video Rotary Position Embedding?
  18. X-Prompt: Generalizable Auto-Regressive Visual Learning with In-Context Prompting
  19. Alpha-CLIP: A CLIP Model Focusing on Wherever you Want
  20. Are We on the Right Way for Evaluating Large Vision-Language Models?
  21. FreeDrag: Feature Dragging for Reliable Point-Based Image Editing
  22. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
  23. Long-CLIP: Unlocking the Long-Text Capability of CLIP
  24. MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
  25. MMLONGBENCH-DOC: Benchmarking Long-context Document Understanding with Visualizations
  26. OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
  27. ShareGPT4V: Improving Large Multi-modal Models with Better Captions
  28. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
  29. Streaming Long Video Understanding with Large Language Models
  30. VIGC: Visual Instruction Generation and Correction
  31. VLMEvalKit: An Open-Source ToolKit for Evaluating Large Multi-Modality Models