PPaperPicks

Shihan Dou

30 papers at tracked venues · 26 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
  2. DARM: Distribution-Aware Reward Modeling by Alleviating Biases from Low Preference-Context Dependency Data
  3. Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
  4. From Scores to Preferences: Redefining Evaluation Paradigm for Speech Quality Reward Modeling
  5. JanusMM: A Benchmark for Self-Deprecation Understanding in Real-World Multimodal Conversations
  6. LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
  7. MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
  8. MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning
  9. Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control
  10. OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
  11. PRISM: Probabilistic Reward Model with Inherent Structural Modeling
  12. VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
  13. Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning
    AAAI 2025 · Shihan Dou
  14. DocFusion: A Unified Framework for Document Parsing Tasks
  15. EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
    NeurIPS 2025 · Shihan Dou
  16. Governance in Motion: Co-evolution of Constitutions and AI models for Scalable Safety
  17. LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
  18. Lost in the Context: Insufficient and Distracted Attention to Contexts in Preference Modeling
    ACL 2025 · Shihan Dou
  19. Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric
  20. Multi-Programming Language Sandbox for LLMs
    ACL 2025 · Shihan Dou
  21. PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts
  22. Pre-Trained Policy Discriminators are General Reward Models
    NeurIPS 2025 · Shihan Dou
  23. RMB: Comprehensively benchmarking reward models in LLM alignment
  24. UPLex: Fine-Grained Personality Control in Large Language Models via Unsupervised Lexical Modulation
  25. Improving Generalization of Alignment with Human Preferences through Group Invariant Learning
  26. Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback
  27. LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
    ACL 2024 · Shihan Dou
  28. StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback
    ACL 2024 · Shihan Dou
  29. Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
  30. TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer Capabilities