PPaperPicks

Haitao Mi

32 papers at tracked venues · 27 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse Domains
  2. EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
  3. Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
  4. Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
  5. Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
  6. Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data
  7. Verified Critical Step Optimization for LLM Agents
  8. WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models
  9. WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
  10. Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience
  11. DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
  12. Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models
  13. Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
  14. Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
  15. Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
  16. LiteSearch: Efficient Tree Search with Dynamic Exploration Budget for Math Reasoning
  17. Low-Bit Quantization Favors Undertrained LLMs
  18. MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation
  19. Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
  20. Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching
  21. The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
  22. Thoughts Are All Over the Place: On the Underthinking of Long Reasoning Models
  23. Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
  24. Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
  25. UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
  26. WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollback
  27. WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World Model
  28. Improving LLM Generations via Fine-Grained Self-Endorsement
  29. Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
  30. Self-Consistency Boosts Calibration for Math Reasoning
  31. The Trickle-down Impact of Reward Inconsistency on RLHF
  32. Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing