PPaperPicks

Philip Torr

University of Oxford, Department of Engineering Science, Oxford, UK

97 papers at tracked venues · 73 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. An MRP Formulation for Supervised Learning: Generalized Temporal Difference Learning Models (Abstract Reprint)
  2. Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation
  3. Can Editing LLMs Inject Harm?
  4. Deep Research Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
  5. Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
  6. AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
  7. Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
  8. Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
  9. CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents
  10. Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
  11. Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
  12. Can an Individual Manipulate the Collective Decisions of Multi-Agents?
  13. Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions
  14. Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
  15. Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models
  16. FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models
  17. Flex3D: Feed-Forward 3D Generation with Flexible Reconstruction Model and Input View Curation
  18. Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
  19. Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
  20. LLM Jailbreak Detection for (Almost) Free!
  21. Large Language Models Miss the Multi-agent Mark
  22. Localizing Events in Videos with Multimodal Queries
  23. MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents
  24. Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
  25. MatchDiffusion: Training-Free Generation of Match-Cuts
  26. Measuring what Matters: Construct Validity in Large Language Model Benchmarks
  27. Minimalist Concept Erasure in Generative Models
  28. Mixture of Experts Made Intrinsically Interpretable
  29. Multimodal Pragmatic Jailbreak on Text-to-image Models
  30. Olympus: A Universal Task Router for Computer Vision Tasks
  31. On the Coexistence and Ensembling of Watermarks
  32. Open-World Objectness Modeling Unifies Novel Object Detection
  33. PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild
  34. Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
  35. PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference Data
  36. Reimagining Safety Alignment with An Image
  37. Shh, don't say that! Domain Certification in LLMs
  38. Stable Virtual Camera: Generative View Synthesis with Diffusion Models
  39. Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
  40. Towards Certification of Uncertainty Calibration under Adversarial Attacks
  41. Towards Interpreting Visual Information Processing in Vision-Language Models
  42. Towards Reliable Identification of Diffusion-based Image Manipulations
  43. VIKI‑R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning
  44. VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
  45. Video Motion Transfer with Diffusion Transformers
  46. Vision-Language Models Do Not Understand Negation
  47. An Image Is Worth 1000 Lies: Transferability of Adversarial Images across Prompts on Vision-Language Models
  48. As Firm As Their Foundations: Creating Transferable Adversarial Examples Across Downstream Tasks with CLIP
  49. CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
  50. CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor
  51. Can Large Language Model Agents Simulate Human Trust Behavior?
  52. Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation
  53. DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
  54. Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer
  55. Efficient Error Certification for Physics-Informed Neural Networks
  56. Efficient Lifelong Model Evaluation in an Era of Rapid Progress
  57. Extracting Training Data From Document-Based VQA Models
  58. FedMedICL: Towards Holistic Evaluation of Distribution Shifts in Federated Medical Imaging
  59. GaussCtrl: Multi-view Consistent Text-Driven 3D Gaussian Splatting Editing
  60. HelloFresh: LLM Evalutions on Streams of Real-World Human Editorial Actions across X Community Notes and Wikipedia edits
  61. Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
  62. Illusory Attacks: Information-theoretic detectability matters in adversarial attacks
  63. Improving Adversarial Transferability via Model Alignment
  64. Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
  65. Influencer Backdoor Attack on Semantic Segmentation
  66. Interpreting Learned Feedback Patterns in Large Language Models
  67. Label Delay in Online Continual Learning
  68. Latent Guard: A Safety Framework for Text-to-Image Generation
  69. Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
  70. NeRF-VPT: Learning Novel View Representations with Neural Radiance Fields via View Prompt Tuning
  71. No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
  72. No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
  73. Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
  74. On Pretraining Data Diversity for Self-Supervised Learning
    ECCV 2024 ·
    Hasan Abed Al Kader Hammoud
  75. PVUW 2024 Challenge on Complex Video Understanding: Methods and Results
  76. Placing Objects in Context via Inpainting for Out-of-Distribution Segmentation
  77. Porf: Pose residual field for accurate Neural surface Reconstruction
  78. Position: Near to Mid-term Risks and Opportunities of Open-Source Generative AI
  79. Prompting a Pretrained Transformer Can Be a Universal Approximator
  80. RanDumb: Random Representations Outperform Online Continually Learned Representations
  81. Real-Fake: Effective Training Data Synthesis Through Distribution Matching
  82. RoDLA: Benchmarking the Robustness of Document Layout Analysis Models
  83. Scene-Conditional 3D Object Stylization and Composition
  84. Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
  85. Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation
  86. Select to Perfect: Imitating desired behavior from large multi-agent data
  87. Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
  88. Set-based Neural Network Encoding Without Weight Tying
  89. Towards Interpretable Deep Local Learning with Successive Gradient Reconciliation
  90. Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
  91. Universal In-Context Approximation By Prompting Fully Recurrent Models
  92. VFusion3D: Learning Scalable 3D Generative Models from Video Diffusion Models
  93. What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
  94. When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
  95. Which Model Generated This Image? A Model-Agnostic Approach for Origin Attribution
  96. WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
  97. uCAP: An Unsupervised Prompting Method for Vision-Language Models