PPaperPicks

Longteng Guo

17 papers at tracked venues · 13 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
    ACL 2026 · Longteng Guo
  2. M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
  3. SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation
    ACL 2026 · Longteng Guo
  4. UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human Trajectories
  5. Ada-K Routing: Boosting the Efficiency of MoE-based LLMs
  6. Breaking the Encoder Barrier for Seamless Video-Language Understanding
  7. Efficient Motion-Aware Video MLLM
  8. GroundingMate: Aiding Object Grounding for Goal-Oriented Vision-and-Language Navigation
  9. Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
  10. VRoPE: Rotary Position Embedding for Video Large Language Models
  11. ViPE: Visual Perception in Parameter Space for Efficient Video-Language Understanding
  12. Collaborative Training of Tiny-Large Vision Language Models
  13. EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
  14. MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
  15. SC- Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models
  16. Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
  17. Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression Segmentation