PPaperPicks

Barbara Plank

LMU Munich, Germany

48 papers at tracked venues · 23 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
  2. Controlling Reading Ease with Gaze-Guided Text Generation
  3. Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
  4. EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
  5. If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
  6. Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
  7. Position: From Noise to Signal to Selbstzweck - Reframing Human Label Variation in the Era of Post-training in NLP
  8. Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
  9. Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models
  10. Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
  11. When Meanings Meet: Investigating the Emergence and Quality of Shared Concept Spaces during Multilingual Language Model Training
  12. ltzGLUE: Luxembourgish General Language Understanding Evaluation
  13. A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
  14. A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI
  15. Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study
  16. Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
  17. Crossing Domains without Labels: Distant Supervision for Term Extraction
  18. Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
  19. Disentangling Subjectivity and Uncertainty for Hate Speech Annotation and Modeling using Gaze
  20. Evaluating Large Language Models for Cross-Lingual Retrieval
  21. LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
  22. LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
  23. Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
  24. M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment Analysis
  25. MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
  26. Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora
  27. Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies Between Model Predictions and Human Responses in VQA
  28. Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges
  29. Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
  30. RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
  31. Reason to Rote: Rethinking Memorization in Reasoning
  32. Refusal Direction is Universal Across Safety-Aligned Languages
  33. Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
  34. The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
  35. Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
  36. Tracing Multilingual Factual Knowledge Acquisition in Pretraining
  37. What Media Frames Reveal About Stance: A Dataset and Study about Memes in Climate Change Discourse
  38. What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
  39. "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
  40. "Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
  41. Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
  42. Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
  43. Position: Insights from Survey Methodology can Improve Training Data
  44. The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models
  45. Through the Lens of Split Vote: Exploring Disagreement, Difficulty and Calibration in Legal Case Outcome Classification
  46. To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity
  47. Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark
  48. VariErr NLI: Separating Annotation Error from Human Label Variation