PPaperPicks

Zhiheng Xi

44 papers at tracked venues · 32 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
  2. AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
    ACL 2026 · Zhiheng Xi
  3. AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
    WWW 2026 · Zhiheng Xi
  4. AgentV-RL: Scaling Reward Modeling with Agentic Verifier
  5. Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
  6. Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
  7. Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
  8. Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
  9. From Scores to Preferences: Redefining Evaluation Paradigm for Speech Quality Reward Modeling
  10. LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
  11. Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
  12. MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
  13. MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning
    AAAI 2026 · Zhiheng Xi
  14. Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
  15. VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
  16. What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
  17. Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment
  18. AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments
    ACL 2025 · Zhiheng Xi
  19. Are LLMs Rational Investors? A Study on the Financial Bias in LLMs
  20. BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
    NeurIPS 2025 · Zhiheng Xi
  21. Better Process Supervision with Bi-directional Rewarding Signals
  22. CritiQ: Mining Data Quality Criteria from Human Preferences
  23. Distill Visual Chart Reasoning Ability from LLMs to MLLMs
  24. Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
  25. LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
  26. LoRACoE: Improving Large Language Model via Composition-based LoRA Expert
  27. Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
  28. Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
  29. Multi-Programming Language Sandbox for LLMs
  30. PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts
  31. Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
  32. Pre-Trained Policy Discriminators are General Reward Models
  33. RMB: Comprehensively benchmarking reward models in LLM alignment
  34. TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use
  35. ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
  36. Toward Optimal LLM Alignments Using Two-Player Games
  37. Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning
  38. Improving Generalization of Alignment with Human Preferences through Group Invariant Learning
  39. Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data
  40. LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
  41. Reward Modeling Requires Automatic Adjustment Based on Data Quality
  42. Self-Demos: Eliciting Out-of-Demonstration Generalizability in Large Language Models
  43. StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback
  44. Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
    ICML 2024 · Zhiheng Xi