PPaperPicks

Dong Yan

8 papers at tracked venues · 7 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time
    ACL 2026 · Dong Yan
  2. 3D-Properties: Identifying Challenges in DPO and Charting a Path Forward
  3. Learning LLM-as-a-Judge for Preference Alignment
  4. Reward Generalization in RLHF: A Topological Perspective
  5. STAIR: Improving Safety Alignment with Introspective Reasoning
  6. Sequential Preference Optimization: Multi-Dimensional Preference Alignment with Implicit Reward Modeling
  7. Exploring the LLM Journey from Cognition to Expression with Linear Representations
  8. Reward Modeling Requires Automatic Adjustment Based on Data Quality