PPaperPicks

Andrew Zisserman

University of Oxford, UK

31 papers at tracked venues · 23 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Open-World Object Counting in Videos
  2. Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing
  3. WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata
  4. A Simple Modality-Agnostic Representation for Scoliosis Phenotyping
  5. Character-Centric Understanding of Animated Movies
  6. Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
  7. From Panels to Prose: Generating Literary Narratives from Comics
  8. LayerLock: Non-Collapsing Representation Learning with Progressive Freezing
  9. Learning from Streaming Video with Orthogonal Gradients
  10. Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues
  11. SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
  12. Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
  13. Understanding Co-Speech Gestures in-the-Wild
  14. 3D Spine Shape Estimation from Single 2D DXA
  15. A General Protocol to Probe Large Vision Models for 3D Physical Understanding
  16. A Simple Recipe for Contrastively Pre-Training Video-First Encoders Beyond 16 Frames
  17. Amodal Ground Truth and Completion in the Wild
  18. Appearance-Based Refinement for Object-Centric Motion Segmentation
  19. AutoAD III: The Prequel - Back to the Pixels
  20. Automated Spinal MRI Labelling from Reports Using a Large Language Model
  21. CountGD: Multi-Modal Open-World Counting
  22. FlexCap: Describe Anything in Images in Controllable Detail
  23. Learning from One Continuous Video Stream
  24. Made to Order: Discovering Monotonic Temporal Changes via Self-supervised Video Ordering
  25. N2F2: Hierarchical Scene Understanding with Nested Neural Feature Fields
  26. Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language
  27. Speech Recognition Models are Strong Lip-readers
  28. TAPVid-3D: A Benchmark for Tracking Any Point in 3D
  29. TIM: A Time Interval Machine for Audio-Visual Action Recognition
  30. Text-Conditioned Resampler For Long Form Video Understanding
  31. The Manga Whisperer: Automatically Generating Transcriptions for Comics