PPaperPicks

Manling Li

23 papers at tracked venues · 20 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. EMCompress: Video-LLMs with Endomorphic Multimodal Compression
  2. Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents
  3. Unifying Inference-Time Planning Language Generation
  4. WorldAgen: Unified State-Action Prediction with Test-Time World Model Training
  5. Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
  6. Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models
  7. EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
  8. Exploring Diffusion Transformer Designs via Grafting
  9. From Large Language Models to Large Action Models: Reasoning and Planning with Physical World Knowledge
    AAAI 2025 · Manling Li
  10. LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World
  11. LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
  12. Re-thinking Temporal Search for Long-Form Video Understanding
  13. SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
  14. The Law of Knowledge Overshadowing: Towards Understanding, Predicting and Preventing LLM Hallucination
  15. VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
  16. Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
  17. Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
    NeurIPS 2024 · Manling Li
  18. HourVideo: 1-Hour Video-Language Understanding
  19. IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
  20. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: EMNLP 2024 - System Demonstrations, Miami, Florida, USA, November 12-16, 2024
    EMNLP 2024 ·
    Delia Irazú Hernández Farías
  21. Training-free Deep Concept Injection Enables Language Models for Video Question Answering
  22. Why Does New Knowledge Create Messy Ripple Effects in LLMs?
  23. Word Embeddings Are Steers for Language Models