P
PaperPicks
Conferences
Peter Henderson
15 papers at tracked venues · 14 at CORE A* · active 2024–2025
DBLP profile ↗
ORCID search ↗
Venues
ICLR
×5
ICML
×4
NeurIPS
×4
AAAI
×1
NAACL
×1
Frequent coauthors
Xiangyu Qi
DBLP profile ↗
ORCID search ↗
×4
Boyi Wei
DBLP profile ↗
ORCID search ↗
×3
Shayne Longpre
DBLP profile ↗
ORCID search ↗
×2
Gaku Morio
DBLP profile ↗
ORCID search ↗
×1
Luxi He
DBLP profile ↗
ORCID search ↗
×1
Joel Niklaus
DBLP profile ↗
ORCID search ↗
×1
Zihan Zheng
DBLP profile ↗
ORCID search ↗
×1
Tinghao Xie
DBLP profile ↗
ORCID search ↗
×1
Sayash Kapoor
DBLP profile ↗
ORCID search ↗
×1
Papers
A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection
NeurIPS 2025
·
Gaku Morio
DBLP profile ↗
ORCID search ↗
Dynamic Risk Assessments for Offensive Cybersecurity Agents
NeurIPS 2025
·
Boyi Wei
DBLP profile ↗
ORCID search ↗
Fantastic Copyrighted Beasts and How (Not) to Generate Them
ICLR 2025
·
Luxi He
DBLP profile ↗
ORCID search ↗
LawInstruct: A Resource for Studying Language Model Adaptation to the Legal Domain
NAACL 2025
·
Joel Niklaus
DBLP profile ↗
ORCID search ↗
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
NeurIPS 2025
·
Zihan Zheng
DBLP profile ↗
ORCID search ↗
On Evaluating the Durability of Safeguards for Open-Weight LLMs
ICLR 2025
·
Xiangyu Qi
DBLP profile ↗
ORCID search ↗
Position: In-House Evaluation Is Not Enough. Towards Robust Third-Party Evaluation and Flaw Disclosure for General-Purpose AI
ICML 2025
·
Shayne Longpre
DBLP profile ↗
ORCID search ↗
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
ICLR 2025
·
Tinghao Xie
DBLP profile ↗
ORCID search ↗
Safety Alignment Should be Made More Than Just a Few Tokens Deep
ICLR 2025
·
Xiangyu Qi
DBLP profile ↗
ORCID search ↗
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
ICML 2024
·
Boyi Wei
DBLP profile ↗
ORCID search ↗
Evaluating Copyright Takedown Methods for Language Models
NeurIPS 2024
·
Boyi Wei
DBLP profile ↗
ORCID search ↗
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
ICLR 2024
·
Xiangyu Qi
DBLP profile ↗
ORCID search ↗
Position: A Safe Harbor for AI Evaluation and Red Teaming
ICML 2024
·
Shayne Longpre
DBLP profile ↗
ORCID search ↗
Position: On the Societal Impact of Open Foundation Models
ICML 2024
·
Sayash Kapoor
DBLP profile ↗
ORCID search ↗
Visual Adversarial Examples Jailbreak Aligned Large Language Models
AAAI 2024
·
Xiangyu Qi
DBLP profile ↗
ORCID search ↗