P
PaperPicks
Conferences
Tao Gui
Fudan University, Shanghai, China
76 papers at tracked venues · 56 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0000-0002-0059-0210 ↗
Google Scholar ↗
Homepage ↗
Venues
ACL
×39
EMNLP
×18
AAAI
×5
NeurIPS
×5
ICLR
×3
NAACL
×2
CVPR
×1
ICCV
×1
ICML
×1
WWW
×1
Frequent coauthors
Junjie Ye
DBLP profile ↗
ORCID search ↗
×7
Shihan Dou
DBLP profile ↗
ORCID search ↗
×7
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
×6
Ming Zhang
DBLP profile ↗
ORCID search ↗
×4
Changhao Jiang
DBLP profile ↗
ORCID search ↗
×2
Xiaoran Fan
DBLP profile ↗
ORCID search ↗
×2
Binghai Wang
DBLP profile ↗
ORCID search ↗
×2
Yuming Yang
DBLP profile ↗
ORCID search ↗
×2
Wei He
DBLP profile ↗
ORCID search ↗
×2
Shuo Li
DBLP profile ↗
ORCID search ↗
×2
Rui Zheng
DBLP profile ↗
ORCID search ↗
×2
Jun Zhao
DBLP profile ↗
ORCID search ↗
×2
Papers
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
ACL 2026
·
Jie Yang
DBLP profile ↗
ORCID search ↗
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
ACL 2026
·
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
WWW 2026
·
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
ACL 2026
·
Jiazheng Zhang
DBLP profile ↗
ORCID search ↗
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
ACL 2026
·
Changhao Jiang
DBLP profile ↗
ORCID search ↗
Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
ACL 2026
·
Xin Guo
DBLP profile ↗
ORCID search ↗
DARM: Distribution-Aware Reward Modeling by Alleviating Biases from Low Preference-Context Dependency Data
ACL 2026
·
Shaofan Liu
DBLP profile ↗
ORCID search ↗
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
ACL 2026
·
Junzhe Wang
DBLP profile ↗
ORCID search ↗
Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
ACL 2026
·
Junjie Ye
DBLP profile ↗
ORCID search ↗
FinToolSyn: A forward synthesis Framework for Financial Tool-Use Dialogue Data with Dynamic Tool Retrieval
ACL 2026
·
Caishuang Huang
DBLP profile ↗
ORCID search ↗
From Scores to Preferences: Redefining Evaluation Paradigm for Speech Quality Reward Modeling
ACL 2026
·
Yifei Cao
DBLP profile ↗
ORCID search ↗
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
ACL 2026
·
Ming Zhang
DBLP profile ↗
ORCID search ↗
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
ACL 2026
·
Hengyuan Zhang
DBLP profile ↗
ORCID search ↗
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention Across Vision-Language Models
AAAI 2026
·
Xiaoran Fan
DBLP profile ↗
ORCID search ↗
MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
ACL 2026
·
Jiahang Lin
DBLP profile ↗
ORCID search ↗
MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning
AAAI 2026
·
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
ACL 2026
·
Junjie Ye
DBLP profile ↗
ORCID search ↗
Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control
ACL 2026
·
Changhao Jiang
DBLP profile ↗
ORCID search ↗
OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
ACL 2026
·
Deming Ding
DBLP profile ↗
ORCID search ↗
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
ACL 2026
·
Binghai Wang
DBLP profile ↗
ORCID search ↗
VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
ACL 2026
·
Dingwei Zhu
DBLP profile ↗
ORCID search ↗
What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
AAAI 2026
·
Xiaoran Fan
DBLP profile ↗
ORCID search ↗
Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment
ACL 2026
·
Yuming Yang
DBLP profile ↗
ORCID search ↗
AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments
ACL 2025
·
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning
AAAI 2025
·
Shihan Dou
DBLP profile ↗
ORCID search ↗
Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
EMNLP 2025
·
Junjie Ye
DBLP profile ↗
ORCID search ↗
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
NeurIPS 2025
·
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
Better Process Supervision with Bi-directional Rewarding Signals
ACL 2025
·
Wenxiang Chen
DBLP profile ↗
ORCID search ↗
CritiQ: Mining Data Quality Criteria from Human Preferences
ACL 2025
·
Honglin Guo
DBLP profile ↗
ORCID search ↗
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
EMNLP 2025
·
Wei He
DBLP profile ↗
ORCID search ↗
EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
NeurIPS 2025
·
Shihan Dou
DBLP profile ↗
ORCID search ↗
Governance in Motion: Co-evolution of Constitutions and AI models for Scalable Safety
EMNLP 2025
·
Chenhao Huang
DBLP profile ↗
ORCID search ↗
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
ICLR 2025
·
Shuo Li
DBLP profile ↗
ORCID search ↗
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
NeurIPS 2025
·
Wujian Peng
DBLP profile ↗
ORCID search ↗
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
EMNLP 2025
·
Ming Zhang
DBLP profile ↗
ORCID search ↗
LoRACoE: Improving Large Language Model via Composition-based LoRA Expert
EMNLP 2025
·
Guanyu Li
DBLP profile ↗
ORCID search ↗
Lost in the Context: Insufficient and Distracted Attention to Contexts in Preference Modeling
ACL 2025
·
Shihan Dou
DBLP profile ↗
ORCID search ↗
Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric
ACL 2025
·
Yuming Yang
DBLP profile ↗
ORCID search ↗
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
EMNLP 2025
·
Shuo Li
DBLP profile ↗
ORCID search ↗
Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
NAACL 2025
·
Yiwen Ding
DBLP profile ↗
ORCID search ↗
Multi-Programming Language Sandbox for LLMs
ACL 2025
·
Shihan Dou
DBLP profile ↗
ORCID search ↗
PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts
ACL 2025
·
Ming Zhang
DBLP profile ↗
ORCID search ↗
Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
EMNLP 2025
·
Senjie Jin
DBLP profile ↗
ORCID search ↗
Pre-Trained Policy Discriminators are General Reward Models
NeurIPS 2025
·
Shihan Dou
DBLP profile ↗
ORCID search ↗
RMB: Comprehensively benchmarking reward models in LLM alignment
ICLR 2025
·
Enyu Zhou
DBLP profile ↗
ORCID search ↗
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Models
CVPR 2025
·
Yongting Zhang
DBLP profile ↗
ORCID search ↗
TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use
EMNLP 2025
·
Junjie Ye
DBLP profile ↗
ORCID search ↗
ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
ACL 2025
·
Junjie Ye
DBLP profile ↗
ORCID search ↗
Toward Optimal LLM Alignments Using Two-Player Games
EMNLP 2025
·
Rui Zheng
DBLP profile ↗
ORCID search ↗
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
ACL 2025
·
Tao Ji
DBLP profile ↗
ORCID search ↗
Understanding Parametric and Contextual Knowledge Reconciliation within Large Language Models
NeurIPS 2025
·
Jun Zhao
DBLP profile ↗
ORCID search ↗
UniPaint: Unified Space-Time Video Inpainting via Mixture-of-Experts
ICCV 2025
·
Zhen Wan
DBLP profile ↗
ORCID search ↗
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
ACL 2024
·
Jun Zhan
DBLP profile ↗
ORCID search ↗
Enhancing Contrastive Learning with Noise-Guided Attack: Towards Continual Relation Extraction in the Wild
ACL 2024
·
Ting Wu
DBLP profile ↗
ORCID search ↗
Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning
EMNLP 2024
·
Lu Chen
DBLP profile ↗
ORCID search ↗
Improving Generalization of Alignment with Human Preferences through Group Invariant Learning
ICLR 2024
·
Rui Zheng
DBLP profile ↗
ORCID search ↗
LLMEval: A Preliminary Study on How to Evaluate Large Language Models
AAAI 2024
·
Yue Zhang
DBLP profile ↗
ORCID search ↗
LONGAGENT: Achieving Question Answering for 128k-Token-Long Documents through Multi-Agent Collaboration
EMNLP 2024
·
Jun Zhao
DBLP profile ↗
ORCID search ↗
Length Generalization of Causal Transformers without Position Encoding
ACL 2024
·
Jie Wang
DBLP profile ↗
ORCID search ↗
LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
ACL 2024
·
Shihan Dou
DBLP profile ↗
ORCID search ↗
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
EMNLP 2024
·
Yi Lu
DBLP profile ↗
ORCID search ↗
Making Harmful Behaviors Unlearnable for Large Language Models
ACL 2024
·
Xin Zhou
DBLP profile ↗
ORCID search ↗
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
EMNLP 2024
·
Chong Zhang
DBLP profile ↗
ORCID search ↗
Navigating the OverKill in Large Language Models
ACL 2024
·
Chenyu Shi
DBLP profile ↗
ORCID search ↗
P4: Plug-and-Play Discrete Prompting for Large Language Models Personalization
ACL 2024
·
Yuansen Zhang
DBLP profile ↗
ORCID search ↗
PDF-to-Tree: Parsing PDF Text Blocks into a Tree
EMNLP 2024
·
Yue Zhang
DBLP profile ↗
ORCID search ↗
Rescue: Ranking LLM Responses with Partial Ordering to Improve Response Generation
ACL 2024
·
Yikun Wang
DBLP profile ↗
ORCID search ↗
Reward Modeling Requires Automatic Adjustment Based on Data Quality
EMNLP 2024
·
Binghai Wang
DBLP profile ↗
ORCID search ↗
RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning
EMNLP 2024
·
Junjie Ye
DBLP profile ↗
ORCID search ↗
Self-Demos: Eliciting Out-of-Demonstration Generalizability in Large Language Models
NAACL 2024
·
Wei He
DBLP profile ↗
ORCID search ↗
StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback
ACL 2024
·
Shihan Dou
DBLP profile ↗
ORCID search ↗
ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages
ACL 2024
·
Junjie Ye
DBLP profile ↗
ORCID search ↗
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
ICML 2024
·
Zhiheng Xi
DBLP profile ↗
ORCID search ↗
TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer Capabilities
EMNLP 2024
·
Ming Zhang
DBLP profile ↗
ORCID search ↗
Unveiling Linguistic Regions in Large Language Models
ACL 2024
·
Zhihao Zhang
DBLP profile ↗
ORCID search ↗
Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs
EMNLP 2024
·
Xin Zhou
DBLP profile ↗
ORCID search ↗