P
PaperPicks
Conferences
Yuhao Zhou
14 papers at tracked venues · 11 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
ACL
×5
EMNLP
×3
AAAI
×2
ICLR
×1
ICML
×1
NeurIPS
×1
WWW
×1
Frequent coauthors
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
×2
Shihan Dou
DBLP profile ↗
ORCID search ↗
×2
Yusong Hu
DBLP profile ↗
ORCID search ↗
×1
Jiahang Lin
DBLP profile ↗
ORCID search ↗
×1
Xiangyu Zhao
DBLP profile ↗
ORCID search ↗
×1
Mingqi Wu
DBLP profile ↗
ORCID search ↗
×1
Xiaoran Fan
DBLP profile ↗
ORCID search ↗
×1
Senjie Jin
DBLP profile ↗
ORCID search ↗
×1
Lu Chen
DBLP profile ↗
ORCID search ↗
×1
Rui Zheng
DBLP profile ↗
ORCID search ↗
×1
Binghai Wang
DBLP profile ↗
ORCID search ↗
×1
Papers
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
WWW 2026
·
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
FlowSearch: Advancing Deep Research with Dynamic Structured Knowledge Flow
ACL 2026
·
Yusong Hu
DBLP profile ↗
ORCID search ↗
MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
ACL 2026
·
Jiahang Lin
DBLP profile ↗
ORCID search ↗
MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs
ACL 2026
·
Xiangyu Zhao
DBLP profile ↗
ORCID search ↗
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
AAAI 2026
·
Mingqi Wu
DBLP profile ↗
ORCID search ↗
What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
AAAI 2026
·
Xiaoran Fan
DBLP profile ↗
ORCID search ↗
Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
EMNLP 2025
·
Senjie Jin
DBLP profile ↗
ORCID search ↗
Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
NeurIPS 2025
·
Yuhao Zhou
Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning
EMNLP 2024
·
Lu Chen
DBLP profile ↗
ORCID search ↗
Improving Generalization of Alignment with Human Preferences through Group Invariant Learning
ICLR 2024
·
Rui Zheng
DBLP profile ↗
ORCID search ↗
LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
ACL 2024
·
Shihan Dou
DBLP profile ↗
ORCID search ↗
Reward Modeling Requires Automatic Adjustment Based on Data Quality
EMNLP 2024
·
Binghai Wang
DBLP profile ↗
ORCID search ↗
StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback
ACL 2024
·
Shihan Dou
DBLP profile ↗
ORCID search ↗
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
ICML 2024
·
Zhiheng Xi
DBLP profile ↗
ORCID search ↗