PPaperPicks

Linfeng Song

20 papers at tracked venues · 19 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse Domains
  2. EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
  3. Verified Critical Step Optimization for LLM Agents
  4. Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience
  5. Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models
  6. Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
  7. Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
  8. Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
  9. LiteSearch: Efficient Tree Search with Dynamic Exploration Budget for Math Reasoning
  10. MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation
  11. SR-LLM: Rethinking the Structured Representation in Large Language Model
  12. Thoughts Are All Over the Place: On the Underthinking of Long Reasoning Models
  13. Discrete Conditional Diffusion for Reranking in Recommendation
  14. Improving LLM Generations via Fine-Grained Self-Endorsement
  15. Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal
  16. Response Enhanced Semi-supervised Dialogue Query Generation
  17. Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
  18. Self-Consistency Boosts Calibration for Math Reasoning
  19. The Trickle-down Impact of Reward Inconsistency on RLHF
  20. Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing