PPaperPicks

Cordelia Schmid

INRIA, France

33 papers at tracked venues · 27 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
  2. Dense Video Object Captioning from Disjoint Supervision
  3. FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement
  4. Flexible Frame Selection for Efficient Video Reasoning
  5. FlowNav: Combining Flow Matching and Depth Priors for Efficient Navigation
  6. HORT: Monocular Hand-held Objects Reconstruction with Transformers
  7. InteractVLM: 3D Interaction Reasoning from 2D Foundational Models
  8. Language-Guided Image Tokenization for Generation
  9. Large-Scale Pre-Training for Grounded Video Caption Generation
  10. Minerva: Evaluating Complex Video Reasoning
  11. OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
  12. Prediction-Powered Causal Inferences
  13. Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
  14. Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-Guided 3D Policy
  15. Towards Zero-Shot Multimodal Machine Translation
  16. ViViDex: Learning Vision-Based Dexterous Manipulation from Human Videos
  17. Visual Lexicon: Rich Image Features in Language Space
  18. mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
  19. A Generative Approach for Wikipedia-Scale Visual Entity Recognition
  20. CoVR: Learning Composed Video Retrieval from Web Video Captions
  21. DataDream: Few-Shot Guided Dataset Generation
  22. Dense Optical Tracking: Connecting the Dots
  23. End-to-End Spatio-Temporal Action Localisation with Video Transformers
  24. Learning Correlation Structures for Vision Transformers
  25. MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
  26. Pixel Aligned Language Models
  27. Retrieval-Enhanced Contrastive Vision-Text Models
  28. SUGAR : Pre-training 3D Visual Representations for Robotics
  29. SceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender Code
  30. Smoke and Mirrors in Causal Downstream Tasks
  31. Streaming Dense Video Captioning
  32. Time-, Memory- and Parameter-Efficient Visual Adaptation
  33. Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach