P
PaperPicks
Conferences
Paul Röttger
21 papers at tracked venues · 12 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
ACL
×7
NAACL
×4
EMNLP
×3
EACL
×2
ICLR
×2
AAAI
×1
ICML
×1
NeurIPS
×1
Frequent coauthors
Carolin Holtermann
DBLP profile ↗
ORCID search ↗
×2
Xinpeng Wang
DBLP profile ↗
ORCID search ↗
×2
Ying Ying Lim
DBLP profile ↗
ORCID search ↗
×1
Giuseppe Russo
DBLP profile ↗
ORCID search ↗
×1
Shamsuddeen Hassan Muhammad
DBLP profile ↗
ORCID search ↗
×1
Matthias Orlikowski
DBLP profile ↗
ORCID search ↗
×1
Manuel Tonneau
DBLP profile ↗
ORCID search ↗
×1
Emanuele Moscato
DBLP profile ↗
ORCID search ↗
×1
Pedro Henrique Luz de Araujo
DBLP profile ↗
ORCID search ↗
×1
Yong Cao
DBLP profile ↗
ORCID search ↗
×1
Dominik Meier
DBLP profile ↗
ORCID search ↗
×1
Fabio Pernisi
DBLP profile ↗
ORCID search ↗
×1
Papers
Bias in the East, Bias in the West: A Bilingual Analysis of LLM Political Bias on U.S.- and China-Related Issues
EACL 2026
·
Ying Ying Lim
DBLP profile ↗
ORCID search ↗
The Pluralistic Moral Gap: Understanding Moral Judgment and Value Differences between Humans and Large Language Models
EACL 2026
·
Giuseppe Russo
DBLP profile ↗
ORCID search ↗
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages
NAACL 2025
·
Shamsuddeen Hassan Muhammad
DBLP profile ↗
ORCID search ↗
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
ACL 2025
·
Carolin Holtermann
DBLP profile ↗
ORCID search ↗
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
ACL 2025
·
Matthias Orlikowski
DBLP profile ↗
ORCID search ↗
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter
ACL 2025
·
Manuel Tonneau
DBLP profile ↗
ORCID search ↗
Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them
EMNLP 2025
·
Emanuele Moscato
DBLP profile ↗
ORCID search ↗
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
EMNLP 2025
·
Pedro Henrique Luz de Araujo
DBLP profile ↗
ORCID search ↗
SafetyPrompts: A Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
AAAI 2025
·
Paul Röttger
Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations
NAACL 2025
·
Yong Cao
DBLP profile ↗
ORCID search ↗
Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
ICLR 2025
·
Xinpeng Wang
DBLP profile ↗
ORCID search ↗
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
EMNLP 2025
·
Dominik Meier
DBLP profile ↗
ORCID search ↗
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
ACL 2024
·
Xinpeng Wang
DBLP profile ↗
ORCID search ↗
Compromesso! Italian Many-Shot Jailbreaks undermine the safety of Large Language Models
ACL 2024
·
Fabio Pernisi
DBLP profile ↗
ORCID search ↗
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
ACL 2024
·
Carolin Holtermann
DBLP profile ↗
ORCID search ↗
Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset
NAACL 2024
·
Janis Goldzycher
DBLP profile ↗
ORCID search ↗
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
ACL 2024
·
Paul Röttger
Position: Near to Mid-term Risks and Opportunities of Open-Source Generative AI
ICML 2024
·
Francisco Eiras
DBLP profile ↗
ORCID search ↗
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
ICLR 2024
·
Federico Bianchi
DBLP profile ↗
ORCID search ↗
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
NeurIPS 2024
·
Hannah Rose Kirk
DBLP profile ↗
ORCID search ↗
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
NAACL 2024
·
Paul Röttger