PPaperPicks

Senjie Jin

10 papers at tracked venues · 8 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
  2. Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
  3. MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning
  4. Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
  5. VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
  6. What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
  7. Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
    EMNLP 2025 · Senjie Jin
  8. SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Models
  9. Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning
  10. Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning