PPaperPicks

Paul Röttger

21 papers at tracked venues · 12 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Bias in the East, Bias in the West: A Bilingual Analysis of LLM Political Bias on U.S.- and China-Related Issues
  2. The Pluralistic Moral Gap: Understanding Moral Judgment and Value Differences between Humans and Large Language Models
  3. AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages
    NAACL 2025 ·
    Shamsuddeen Hassan Muhammad
  4. Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
  5. Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
  6. HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter
  7. Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them
  8. Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
    EMNLP 2025 ·
    Pedro Henrique Luz de Araujo
  9. SafetyPrompts: A Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
    AAAI 2025 · Paul Röttger
  10. Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations
  11. Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
  12. TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
  13. "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
  14. Compromesso! Italian Many-Shot Jailbreaks undermine the safety of Large Language Models
  15. Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
  16. Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset
  17. Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
    ACL 2024 · Paul Röttger
  18. Position: Near to Mid-term Risks and Opportunities of Open-Source Generative AI
  19. Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
  20. The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
  21. XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
    NAACL 2024 · Paul Röttger