PPaperPicks

Yongbin Li

Capital Normal University, College of Life Sciences, Beijing, China

47 papers at tracked venues · 39 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Act-Adaptive Margin: Dynamically Calibrating Reward Models for Subjective Ambiguity
  2. Beyond Quantity: Trajectory Diversity Scaling for Code Agents
  3. Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions
  4. EvoRoute: Experience-Driven Self-Routing LLM Agent Systems
  5. ExpSeek: Self-Triggered Experience Seeking for Web Agents
  6. Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format
  7. Large Language Model Unlearning for Source Code
  8. MOA: Multi-Objective Alignment for Role-Playing Agents
  9. MemPO: Self-Memory Policy Optimization for Long-Horizon Agents
  10. RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
  11. Reasoning-Guided Exploration for Online DPO
  12. Selective Weak-to-Strong Generalization
  13. To Diff or Not to Diff? Structure-Aware and Adaptive Output Formats for Efficient LLM-based Code Editing
  14. ToM-Synth: Scaling Robust Theory of Mind in LLMs via 6, 912 Structured Social Units
  15. Understanding Generalization in Role-Playing Models via Information Theory
  16. CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization
  17. DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
  18. Debate Helps Weak-to-Strong Generalization
  19. DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking
  20. EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
  21. EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning
  22. ExploraCoder: Advancing Code Generation for Multiple Unseen APIs via Planning and Chained Exploration
  23. IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
  24. MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
  25. OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
  26. On the Role of Attention Heads in Large Language Model Safety
  27. OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis
  28. Reverse Preference Optimization for Complex Instruction Following
  29. SDPO: Segment-Level Direct Preference Optimization for Social Agents
  30. StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization
  31. Supervised Optimism Correction: Be Confident When LLMs Are Sure
  32. Transferable Post-training via Inverse Value Learning
  33. DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
  34. EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
  35. Fine-Tuning Language Models with Reward Learning on Policy
  36. FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
  37. Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use
  38. How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
  39. Improving Factual Consistency of News Summarization by Contrastive Preference Optimization
  40. Iterative Forward Tuning Boosts In-Context Learning in Language Models
  41. Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
  42. Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
  43. Masked Thought: Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language Models
  44. One-Shot Learning as Instruction Data Prospector for Large Language Models
  45. Preference Ranking Optimization for Human Alignment
  46. Self-Retrieval: End-to-End Information Retrieval with One Large Language Model
  47. SoFA: Shielded On-the-fly Alignment via Priority Rule Following