PPaperPicks

Qi Yi

8 papers at tracked venues · 8 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AT²PO: Agentic Turn-based Policy Optimization via Tree Search
  2. Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
  3. Reinforcement Learning on Pre-Training Data
  4. Rhombus: Incentivizing Coordination in Parallel Thinking through Reinforcement Learning
  5. Emergent Communication for Numerical Concepts Generalization
  6. Hypothesis, Verification, and Induction: Grounding Large Language Models with Self-Driven Skill Learning
  7. OCEAN-MBRL: Offline Conservative Exploration for Model-Based Offline Reinforcement Learning
  8. Prompt-based Visual Alignment for Zero-shot Policy Transfer