P
PaperPicks
Conferences
Shihan Dou
30 papers at tracked venues · 26 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
ACL
×18
EMNLP
×4
AAAI
×2
ICLR
×2
ICML
×2
NeurIPS
×2
Frequent coauthors
Ming Zhang
DBLP profile ↗
ORCID search ↗
×4
Changhao Jiang
DBLP profile ↗
ORCID search ↗
×2
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
×2
Shaofan Liu
DBLP profile ↗
ORCID search ↗
×1
Junzhe Wang
DBLP profile ↗
ORCID search ↗
×1
Yifei Cao
DBLP profile ↗
ORCID search ↗
×1
Xinyi Xu
DBLP profile ↗
ORCID search ↗
×1
Jiahang Lin
DBLP profile ↗
ORCID search ↗
×1
Deming Ding
DBLP profile ↗
ORCID search ↗
×1
Yuhang Zhou
DBLP profile ↗
ORCID search ↗
×1
Dingwei Zhu
DBLP profile ↗
ORCID search ↗
×1
Mingxu Chai
DBLP profile ↗
ORCID search ↗
×1
Papers
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
ACL 2026
·
Changhao Jiang
DBLP profile ↗
ORCID search ↗
DARM: Distribution-Aware Reward Modeling by Alleviating Biases from Low Preference-Context Dependency Data
ACL 2026
·
Shaofan Liu
DBLP profile ↗
ORCID search ↗
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
ACL 2026
·
Junzhe Wang
DBLP profile ↗
ORCID search ↗
From Scores to Preferences: Redefining Evaluation Paradigm for Speech Quality Reward Modeling
ACL 2026
·
Yifei Cao
DBLP profile ↗
ORCID search ↗
JanusMM: A Benchmark for Self-Deprecation Understanding in Real-World Multimodal Conversations
ACL 2026
·
Xinyi Xu
DBLP profile ↗
ORCID search ↗
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
ACL 2026
·
Ming Zhang
DBLP profile ↗
ORCID search ↗
MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
ACL 2026
·
Jiahang Lin
DBLP profile ↗
ORCID search ↗
MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning
AAAI 2026
·
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control
ACL 2026
·
Changhao Jiang
DBLP profile ↗
ORCID search ↗
OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
ACL 2026
·
Deming Ding
DBLP profile ↗
ORCID search ↗
PRISM: Probabilistic Reward Model with Inherent Structural Modeling
ACL 2026
·
Yuhang Zhou
DBLP profile ↗
ORCID search ↗
VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
ACL 2026
·
Dingwei Zhu
DBLP profile ↗
ORCID search ↗
Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning
AAAI 2025
·
Shihan Dou
DocFusion: A Unified Framework for Document Parsing Tasks
ACL 2025
·
Mingxu Chai
DBLP profile ↗
ORCID search ↗
EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
NeurIPS 2025
·
Shihan Dou
Governance in Motion: Co-evolution of Constitutions and AI models for Scalable Safety
EMNLP 2025
·
Chenhao Huang
DBLP profile ↗
ORCID search ↗
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
EMNLP 2025
·
Ming Zhang
DBLP profile ↗
ORCID search ↗
Lost in the Context: Insufficient and Distracted Attention to Contexts in Preference Modeling
ACL 2025
·
Shihan Dou
Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric
ACL 2025
·
Yuming Yang
DBLP profile ↗
ORCID search ↗
Multi-Programming Language Sandbox for LLMs
ACL 2025
·
Shihan Dou
PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts
ACL 2025
·
Ming Zhang
DBLP profile ↗
ORCID search ↗
Pre-Trained Policy Discriminators are General Reward Models
NeurIPS 2025
·
Shihan Dou
RMB: Comprehensively benchmarking reward models in LLM alignment
ICLR 2025
·
Enyu Zhou
DBLP profile ↗
ORCID search ↗
UPLex: Fine-Grained Personality Control in Large Language Models via Unsupervised Lexical Modulation
EMNLP 2025
·
Tianlong Li
DBLP profile ↗
ORCID search ↗
Improving Generalization of Alignment with Human Preferences through Group Invariant Learning
ICLR 2024
·
Rui Zheng
DBLP profile ↗
ORCID search ↗
Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback
ICML 2024
·
Songyang Gao
DBLP profile ↗
ORCID search ↗
LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
ACL 2024
·
Shihan Dou
StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback
ACL 2024
·
Shihan Dou
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
ICML 2024
·
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer Capabilities
EMNLP 2024
·
Ming Zhang
DBLP profile ↗
ORCID search ↗