PPaperPicks

Chunyuan Li

19 papers at tracked venues · 13 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
  2. Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
  3. Graphic Design with Large Multimodal Model
  4. LLaVA-Critic: Learning to Evaluate Multimodal Models
  5. LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
  6. LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
  7. MMSearch: Unveiling the Potential of Large Models as Multi-modal Search Engines
  8. MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
  9. Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning
  10. Aligning Large Multimodal Models with Factually Augmented RLHF
  11. Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection
  12. Improved Baselines with Visual Instruction Tuning
  13. LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
  14. LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
  15. MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
  16. Position: TrustLLM: Trustworthiness in Large Language Models
  17. Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
  18. Segment and Recognize Anything at Any Granularity
  19. Visual in-Context Prompting