PPaperPicks

Yuanzhao Zhai

8 papers at tracked venues · 8 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. EvoNarrator: Modeling Scientific Evolution for Feasible Hypothesis Generation
  2. COPR: Continual Human Preference Learning via Optimal Policy Regularization
  3. Correcting Large Language Model Behavior via Influence Function
  4. Empowering Large Language Model Agent through Step-Level Self-Critique and Self-Training
    SIGIR 2025 · Yuanzhao Zhai
  5. Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models
    AAAI 2025 · Yuanzhao Zhai
  6. Preference-Strength-Aware Self-Improving Alignment with Generative Preference Models
    SIGIR 2025 · Yuanzhao Zhai
  7. Iterative Regularized Policy Optimization with Imperfect Demonstrations
  8. Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
    AAAI 2024 · Yuanzhao Zhai