PPaperPicks

Zhuoran Jin

24 papers at tracked venues · 17 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning
  2. Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do
    ACL 2026 · Zhuoran Jin
  3. Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation
  4. Towards Explainable Diagnosis: A Self-learned Explanatory Knowledge Base Approach
  5. A Troublemaker with Contagious Jailbreak Makes Chaos in Honest Towns
  6. Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
  7. Beyond Under-Alignment: Atomic Preference Enhanced Factuality Tuning for Large Language Models
  8. CITI: Enhancing Tool Utilizing Ability in Large Language Models Without Sacrificing General Performance
  9. DTELS: Towards Dynamic Granularity of Timeline Summarization
  10. Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis
  11. Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and Proactivity
  12. FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain
  13. MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models
  14. RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
    ACL 2025 · Zhuoran Jin
  15. RULE: Reinforcement UnLEarning Achieves Forget-retain Pareto Optimality
  16. Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
  17. AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation
  18. Cutting Off the Head Ends the Conflict: A Mechanism for Interpreting and Mitigating Knowledge Conflicts in Language Models
    ACL 2024 · Zhuoran Jin
  19. Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
  20. LINKED: Eliciting, Filtering and Integrating Knowledge in Large Language Model for Commonsense Reasoning
  21. MULFE: A Multi-Level Benchmark for Free Text Model Editing
  22. RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language Models
    NeurIPS 2024 · Zhuoran Jin
  23. Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models
  24. Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language Models