PPaperPicks

Katherine Metcalf

8 papers at tracked venues · 7 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. How Value Induction Reshapes LLM Behavior
  2. Aligning LLMs by Predicting Preferences from User Writing Samples
    ICML 2025 ·
    Stéphane Aroca-Ouellette
  3. Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
  4. Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
  5. Can You Rely on Synthetic Labellers in Preference-Based Reinforcement Learning? It's Complicated
    AAAI 2024 · Katherine Metcalf
  6. Hindsight PRIORs for Reward Learning from Human Preferences
  7. On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
  8. Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models