PPaperPicks

Qi Wu

University of Adelaide, School of Computer Science, Australian Centre for Robotic Vision, Adelaide, Australia

40 papers at tracked venues · 28 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. EgoMemory: Memory-Augmented Personalized Retrieval for Long-Context Egocentric Video
  2. MMCLIP: Cross-Modal Attention Masked Modelling for Medical Language-Image Pre-Training
  3. Manipulation Intention Understanding for Zero-Shot Composed Image Retrieval
  4. OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
  5. VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation Agents
  6. Are Large Vision Language Models Good Game Players?
  7. COSMO: Combination of Selective Memorization for Low-Cost Vision-and-Language Navigation
  8. General Scene Adaptation for Vision-and-Language Navigation
  9. Ground-Level Viewpoint Vision-and-Language Navigation in Continuous Environments
  10. GroundingMate: Aiding Object Grounding for Goal-Oriented Vision-and-Language Navigation
  11. MFL-Owner: Ownership Protection for Multi-modal Federated Learning via Orthogonal Transform Watermark
  12. MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge Distillation
  13. Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval
  14. NavBench: Probing Multimodal Large Language Models for Embodied Navigation
  15. Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
  16. Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval
  17. SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
  18. Secure and Efficient Watermarking for Latent Diffusion Models in Model Distribution Scenarios
  19. SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
  20. VLN-ChEnv: Vision-language Navigation in Changeable Environments
  21. Augmented Commonsense Knowledge for Remote Object Grounding
  22. Context-I2W: Mapping Images to Context-Dependent Words for Accurate Zero-Shot Composed Image Retrieval
  23. Continual Self-Supervised Learning: Towards Universal Multi-Modal Medical Data Representation Learning
  24. Dataset, Challenge, and Evaluation for Tumor Segmentation Variability
  25. Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-Training Framework
  26. Everyday Object Meets Vision-and-Language Navigation Agent via Backdoor
  27. G-NeRF: Geometry-enhanced Novel View Synthesis from Single-View Images
  28. Improving Online Source-Free Domain Adaptation for Object Detection by Unsupervised Data Acquisition
  29. LLM as Copilot for Coarse-Grained Vision-and-Language Navigation
  30. ModaVerse: Efficiently Transforming Modalities with LLMs
  31. NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
  32. NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
  33. NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
  34. Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
  35. PairAug: What Can Augmented Image-Text Pairs Do for Radiology?
  36. Spot the Difference: Difference Visual Question Answering with Residual Alignment
  37. T2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval
  38. Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning
  39. WebVLN: Vision-and-Language Navigation on Websites
  40. Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts