PPaperPicks

Dian Yu

15 papers at tracked venues · 11 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse Domains
  2. Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
  3. Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience
  4. DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
  5. Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models
  6. Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
  7. Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
  8. Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
  9. LiteSearch: Efficient Tree Search with Dynamic Exploration Budget for Math Reasoning
  10. Thoughts Are All Over the Place: On the Underthinking of Long Reasoning Models
  11. Alternate Diverse Teaching for Semi-supervised Medical Image Segmentation
  12. Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
  13. Skills-in-Context: Unlocking Compositionality in Large Language Models
  14. Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic Representations
  15. Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing