PPaperPicks

Peter Henderson

15 papers at tracked venues · 14 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection
  2. Dynamic Risk Assessments for Offensive Cybersecurity Agents
  3. Fantastic Copyrighted Beasts and How (Not) to Generate Them
  4. LawInstruct: A Resource for Studying Language Model Adaptation to the Legal Domain
  5. LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
  6. On Evaluating the Durability of Safeguards for Open-Weight LLMs
  7. Position: In-House Evaluation Is Not Enough. Towards Robust Third-Party Evaluation and Flaw Disclosure for General-Purpose AI
  8. SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
  9. Safety Alignment Should be Made More Than Just a Few Tokens Deep
  10. Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
  11. Evaluating Copyright Takedown Methods for Language Models
  12. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
  13. Position: A Safe Harbor for AI Evaluation and Red Teaming
  14. Position: On the Societal Impact of Open Foundation Models
  15. Visual Adversarial Examples Jailbreak Aligned Large Language Models