PPaperPicks

Jing Liu

Chinese Academy of Sciences, Institute of Automation, National Laboratory of Pattern Recognition, Beijing, China

33 papers at tracked venues · 25 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
  2. M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
  3. SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation
  4. UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human Trajectories
  5. AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion
  6. Ada-K Routing: Boosting the Efficiency of MoE-based LLMs
  7. Balancing Precision and Generalization: Dynamic Instruction Generation for Model Adaptive Zero-Shot Reasoning in LLMs
  8. Breaking the Encoder Barrier for Seamless Video-Language Understanding
  9. C-NAV: Towards Self-Evolving Continual Object Navigation in Open World
  10. COSMO: Combination of Selective Memorization for Low-Cost Vision-and-Language Navigation
  11. Diffusion Feedback Helps CLIP See Better
  12. Efficient Motion-Aware Video MLLM
  13. End-to-End Vision Tokenizer Tuning
  14. GroundingMate: Aiding Object Grounding for Goal-Oriented Vision-and-Language Navigation
  15. Learning Beyond Still Frames: Scaling Vision-Language Models with Video
  16. MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge Distillation
  17. Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
  18. Numerical Pruning for Efficient Autoregressive Models
  19. QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge
  20. Scaling Omni-Modal Pretraining with Multimodal Context: Advancing Universal Representation Learning Across Modalities
  21. VRoPE: Rotary Position Embedding for Video Large Language Models
  22. ViPE: Visual Perception in Parameter Space for Efficient Video-Language Understanding
  23. Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions
  24. COSA: Concatenated Sample Pretrained Vision-Language Foundation Model
  25. Collaborative Training of Tiny-Large Vision Language Models
  26. LLM as Copilot for Coarse-Grained Vision-and-Language Navigation
  27. LSVOS Challenge Report: Large-Scale Complex and Long Video Object Segmentation
  28. MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
  29. PVUW 2024 Challenge on Complex Video Understanding: Methods and Results
  30. SC- Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models
  31. Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
  32. Soft Knowledge Prompt: Help External Knowledge Become a Better Teacher to Instruct LLM in Knowledge-based VQA
  33. Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression Segmentation