PPaperPicks

Ran Xu

Salesforce Research, Salesforce AI Research,

18 papers at tracked venues · 11 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. BLIP-3: A Family of Open Large Multimodal Models
  2. Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
  3. DyMU: Dynamic Merging and Virtual Unmerging for Efficient Variable-Length VLMs
  4. Structured Policy Optimization: Enhance Large Vision-Language Model via Self-Referenced Dialogue
  5. Text2Data: Low-Resource Data Generation with Textual Control
  6. Trust but Verify: Programmatic VLM Evaluation in the Wild
  7. xLAM: A Family of Large Action Models to Empower AI Agent Systems
  8. FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
  9. HIVE: Harnessing Human Feedback for Instructional Visual Editing
  10. Hierarchical Point Attention for Indoor 3D Object Detection
  11. LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer
  12. MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
  13. Position: TrustLLM: Trustworthiness in Large Language Models
  14. Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization
  15. SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
  16. ULIP-2: Towards Scalable Multimodal Pre-Training for 3D Understanding
  17. X-InstructBLIP: A Framework for Aligning Image, 3D, Audio, Video to LLMs and its Emergent Cross-Modal Reasoning
  18. xGen-VideoSyn-1: High-Fidelity Text-to-Video Synthesis with Compressed Representations