PPaperPicks

Jiaming Ji

30 papers at tracked venues · 30 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A Game-Theoretica Negotiation Framework for Cross-Cultural Consensus
  2. AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
  3. Benchmarking Fine-Grained Error Detection in Multimodal Reasoning
  4. Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across Modalities
  5. SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning
  6. SafeMT: Multi-turn Safety for Multimodal Language Models
  7. What, Whether and How? Unveiling Process Reward Models for Thinking with Images Reasoning
  8. When Slower Isn't Truer: Inverse Scaling Law of Truthfulness in Multimodal Reasoning
  9. A Survey of LLM-based Agents in Medicine: How far are we from Baymax?
  10. Benchmarking Multi-National Value Alignment for Large Language Models
  11. Boosting Policy and Process Reward Models with Monte Carlo Tree Search in Open-Domain QA
  12. FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation
  13. Generative RLHF-V: Learning Principles from Multi-modal Human Preference
  14. InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
  15. Language Models Resist Alignment: Evidence From Data Compression
    ACL 2025 · Jiaming Ji
  16. LegalReasoner: Step-wised Verification-Correction for Legal Judgment Reasoning
  17. PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
  18. PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
    ACL 2025 · Jiaming Ji
  19. Reward Generalization in RLHF: A Topological Perspective
  20. SAE-V: Interpreting Multimodal Models for Enhanced Alignment
  21. Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
    NeurIPS 2025 · Jiaming Ji
  22. SafeLawBench: Towards Safe Alignment of Large Language Models
  23. SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
  24. Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback
  25. Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction
  26. Aligner: Efficient Alignment by Learning to Correct
    NeurIPS 2024 · Jiaming Ji
  27. ProgressGym: Alignment with a Millennium of Moral Progress
  28. Safe RLHF: Safe Reinforcement Learning from Human Feedback
  29. SafeDreamer: Safe Reinforcement Learning with World Models
  30. SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset