PPaperPicks

Bowen Yu

Alibaba Group

24 papers at tracked venues · 16 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. On the Editability of Delta Parameters in Post-Trained Models
  2. Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
  3. PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
  4. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
  5. Efficient Long Context Fine-tuning with Chunk Flow
  6. LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
  7. MARGE: Improving Math Reasoning with Guided Exploration
  8. P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
  9. ProcessBench: Identifying Process Errors in Mathematical Reasoning
  10. RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing
  11. Rethinking Data Selection at Scale: Random Selection is Almost All You Need
  12. START: Self-taught Reasoner with Tools
  13. Self-Steering Optimization: Autonomous Preference Optimization for Large Language Models
  14. Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
  15. The Lessons of Developing Process Reward Models in Mathematical Reasoning
  16. Transferable Post-training via Inverse Value Learning
  17. Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
  18. Language Models can Evaluate Themselves via Probability Discrepancy
  19. Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment
  20. Predicting Rewards Alongside Tokens: Non-disruptive Parameter Insertion for Efficient Inference Intervention in Large Language Model
  21. Preference Ranking Optimization for Human Alignment
  22. Self-Retrieval: End-to-End Information Retrieval with One Large Language Model
  23. SoFA: Shielded On-the-fly Alignment via Priority Rule Following
  24. TRUE-UIE: Two Universal Relations Unify Information Extraction Tasks