PPaperPicks

Dayiheng Liu

20 papers at tracked venues · 17 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
  2. HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
  3. MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
  4. PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
  5. Chain of Execution Supervision Promotes General Reasoning in Large Language Models
  6. DataMan: Data Manager for Pre-training Large Language Models
  7. Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
  8. Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
  9. HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning
  10. LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
  11. NOVA-63: Native Omni-lingual Versatile Assessments of 63 Disciplines
  12. P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
  13. Parallel Scaling Law for Language Models
  14. ProcessBench: Identifying Process Errors in Mathematical Reasoning
  15. START: Self-taught Reasoner with Tools
  16. Teaching Language Models to Reason with Tools
  17. The Lessons of Developing Process Reward Models in Mathematical Reasoning
  18. How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
  19. Rationales for Answers to Simple Math Word Problems Confuse Large Language Models
  20. Talk Funny! A Large-Scale Humor Response Dataset with Chain-of-Humor Interpretation