P
PaperPicks
Conferences
Xiangyu Qi
10 papers at tracked venues · 9 at CORE A* · active 2024–2025
DBLP profile ↗
ORCID search ↗
Venues
ICLR
×5
AAAI
×1
ACL
×1
ICML
×1
NAACL
×1
NeurIPS
×1
Frequent coauthors
Tinghao Xie
DBLP profile ↗
ORCID search ↗
×2
Chen Xiong
DBLP profile ↗
ORCID search ↗
×1
Haonan Li
DBLP profile ↗
ORCID search ↗
×1
Boyi Wei
DBLP profile ↗
ORCID search ↗
×1
Jiongxiao Wang
DBLP profile ↗
ORCID search ↗
×1
Papers
Defensive Prompt Patch: A Robust and Generalizable Defense of Large Language Models against Jailbreak Attacks
ACL 2025
·
Chen Xiong
DBLP profile ↗
ORCID search ↗
Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability
NAACL 2025
·
Haonan Li
DBLP profile ↗
ORCID search ↗
On Evaluating the Durability of Safeguards for Open-Weight LLMs
ICLR 2025
·
Xiangyu Qi
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
ICLR 2025
·
Tinghao Xie
DBLP profile ↗
ORCID search ↗
Safety Alignment Should be Made More Than Just a Few Tokens Deep
ICLR 2025
·
Xiangyu Qi
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
ICML 2024
·
Boyi Wei
DBLP profile ↗
ORCID search ↗
BaDExpert: Extracting Backdoor Functionality for Accurate Backdoor Input Detection
ICLR 2024
·
Tinghao Xie
DBLP profile ↗
ORCID search ↗
BackdoorAlign: Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
NeurIPS 2024
·
Jiongxiao Wang
DBLP profile ↗
ORCID search ↗
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
ICLR 2024
·
Xiangyu Qi
Visual Adversarial Examples Jailbreak Aligned Large Language Models
AAAI 2024
·
Xiangyu Qi