PPaperPicks

Tao Gui

Fudan University, Shanghai, China

76 papers at tracked venues · 56 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
  2. AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
  3. AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
  4. AgentV-RL: Scaling Reward Modeling with Agentic Verifier
  5. Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
  6. Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
  7. DARM: Distribution-Aware Reward Modeling by Alleviating Biases from Low Preference-Context Dependency Data
  8. Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
  9. Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
  10. FinToolSyn: A forward synthesis Framework for Financial Tool-Use Dialogue Data with Dynamic Tool Retrieval
  11. From Scores to Preferences: Redefining Evaluation Paradigm for Speech Quality Reward Modeling
  12. LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
  13. Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
  14. MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention Across Vision-Language Models
  15. MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
  16. MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning
  17. MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
  18. Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control
  19. OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
  20. Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
  21. VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
  22. What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
  23. Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment
  24. AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments
  25. Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning
  26. Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
  27. BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
  28. Better Process Supervision with Bi-directional Rewarding Signals
  29. CritiQ: Mining Data Quality Criteria from Human Preferences
  30. Distill Visual Chart Reasoning Ability from LLMs to MLLMs
  31. EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
  32. Governance in Motion: Co-evolution of Constitutions and AI models for Scalable Safety
  33. Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
  34. INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
  35. LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
  36. LoRACoE: Improving Large Language Model via Composition-based LoRA Expert
  37. Lost in the Context: Insufficient and Distracted Attention to Contexts in Preference Modeling
  38. Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric
  39. Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
  40. Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
  41. Multi-Programming Language Sandbox for LLMs
  42. PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts
  43. Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
  44. Pre-Trained Policy Discriminators are General Reward Models
  45. RMB: Comprehensively benchmarking reward models in LLM alignment
  46. SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Models
  47. TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use
  48. ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
  49. Toward Optimal LLM Alignments Using Two-Player Games
  50. Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
  51. Understanding Parametric and Contextual Knowledge Reconciliation within Large Language Models
  52. UniPaint: Unified Space-Time Video Inpainting via Mixture-of-Experts
  53. AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
  54. Enhancing Contrastive Learning with Noise-Guided Attack: Towards Continual Relation Extraction in the Wild
  55. Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning
  56. Improving Generalization of Alignment with Human Preferences through Group Invariant Learning
  57. LLMEval: A Preliminary Study on How to Evaluate Large Language Models
  58. LONGAGENT: Achieving Question Answering for 128k-Token-Long Documents through Multi-Agent Collaboration
  59. Length Generalization of Causal Transformers without Position Encoding
  60. LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
  61. LongHeads: Multi-Head Attention is Secretly a Long Context Processor
  62. Making Harmful Behaviors Unlearnable for Large Language Models
  63. Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
  64. Navigating the OverKill in Large Language Models
  65. P4: Plug-and-Play Discrete Prompting for Large Language Models Personalization
  66. PDF-to-Tree: Parsing PDF Text Blocks into a Tree
  67. Rescue: Ranking LLM Responses with Partial Ordering to Improve Response Generation
  68. Reward Modeling Requires Automatic Adjustment Based on Data Quality
  69. RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning
  70. Self-Demos: Eliciting Out-of-Demonstration Generalizability in Large Language Models
  71. StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback
  72. ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages
  73. Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
  74. TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer Capabilities
  75. Unveiling Linguistic Regions in Large Language Models
  76. Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs