PPaperPicks

Adam Gleave

7 papers at tracked venues · 6 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. STACK: Adversarial Attacks on LLM Safeguard Pipelines
  2. Can Go AIs Be Adversarially Robust?
  3. Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
  4. Preference Learning with Lie Detectors can Induce Honesty or Evasion
  5. Scaling Trends for Data Poisoning in LLMs
  6. Scaling Trends in Language Model Robustness
  7. STARC: A General Framework For Quantifying Differences Between Reward Functions