PPaperPicks

Dongrui Liu

23 papers at tracked venues · 21 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
  2. AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent Systems
  3. Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
  4. IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
  5. LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
  6. ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
  7. Cooperative or Competitive? Understanding the Interaction between Attention Heads From A Game Theory Perspective
  8. Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
  9. Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
  10. EvoBench: Towards Real-world LLM-Generated Text Detection Benchmarking for Evolving Large Language Models
  11. LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint
  12. LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
  13. REEF: Representation Encoding Fingerprints for Large Language Models
  14. RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
  15. The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
  16. The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models
  17. VLSBench: Unveiling Visual Leakage in Multimodal Safety
  18. X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Jailbreak Attacks without Compromising Usability
  19. Explaining Generalization Power of a DNN Using Interactive Concepts
  20. Identifying Semantic Induction Heads to Understand In-Context Learning
  21. MLP Can Be a Good Transformer Learner
  22. Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models
  23. Towards the Dynamics of a DNN Learning Symbolic Interactions