PPaperPicks

Rui Zheng

25 papers at tracked venues · 19 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
  2. MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning
  3. Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning Model
  4. What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
  5. AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments
  6. Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning
  7. Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
  8. Multi-Programming Language Sandbox for LLMs
  9. RMB: Comprehensively benchmarking reward models in LLM alignment
  10. SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Models
  11. Toward Optimal LLM Alignments Using Two-Player Games
    EMNLP 2025 · Rui Zheng
  12. Application of nnUnet for Multi-class Segmentation of Aortic Branches and Zones in CTA
  13. DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation
  14. Enhancing Contrastive Learning with Noise-Guided Attack: Towards Continual Relation Extraction in the Wild
  15. Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning
  16. Improving Generalization of Alignment with Human Preferences through Group Invariant Learning
    ICLR 2024 · Rui Zheng
  17. Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback
  18. LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
  19. Reliable Source Approximation: Source-Free Unsupervised Domain Adaptation for Vestibular Schwannoma MRI Segmentation
  20. Rescue: Ranking LLM Responses with Partial Ordering to Improve Response Generation
  21. Reward Modeling Requires Automatic Adjustment Based on Data Quality
  22. RoCoSDF: Row-Column Scanned Neural Signed Distance Fields for Freehand 3D Ultrasound Imaging Shape Reconstruction
  23. StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback
  24. Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
  25. Uncertainty Aware Learning for Language Model Alignment