PPaperPicks

Yunhao Tang

12 papers at tracked venues · 11 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. A Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning
  2. Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
  3. Beyond Verifiable Rewards: Scaling Reinforcement Learning in Language Models to Unverifiable Data
    NeurIPS 2025 · Yunhao Tang
  4. Categorical Distributional Reinforcement Learning with Kullback-Leibler Divergence: Convergence and Asymptotics
  5. Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
    ICML 2025 · Yunhao Tang
  6. A Distributional Analogue to the Successor Representation
  7. Generalized Preference Optimization: A Unified Approach to Offline Alignment
    ICML 2024 · Yunhao Tang
  8. Human Alignment of Large Language Models through Online Preference Optimisation
  9. Learning Uncertainty-Aware Temporally-Extended Actions
  10. Nash Learning from Human Feedback
  11. Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
  12. On scalable oversight with weak LLMs judging strong LLMs