PPaperPicks

Lei Hou

Tsinghua University, Department of Computer Science and Technology, BDCST, BNRist, KIRC, Institute for Artificial Intelligence, Beijing, China

52 papers at tracked venues · 36 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models
  2. Can Large Language Models Effectively Support Decision-Making in Sudden Emergencies?
  3. Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards
  4. DeepPrune: Parallel Scaling without Inter-trace Redundancy
  5. From Knowing to Teaching: Scaffolding Pedagogical Decisions for LLM Agent
  6. Personalized Learning Path Planning through Goal-Driven Learner State Modeling
  7. RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
  8. SimPBL: A Multi-Agent Framework for Project-Based Learning
  9. WildReward: Learning Reward Models from In-the-Wild Human Interactions
  10. AGENTIF: Benchmarking Large Language Models Instruction Following Ability in Agentic Scenarios
  11. AdaptThink: Reasoning Models Can Learn When to Think
  12. Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
  13. AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning
  14. Awaking the Slides: A Tuning-free and Knowledge-regulated AI Tutoring System via Language Model Coordination
  15. CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
  16. Constraint Back-translation Improves Complex Instruction Following of Large Language Models
  17. EduCraft: A System for Generating Pedagogical Lecture Scripts from Long-Context Multimodal Presentations
  18. Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis
  19. EventSum: A Large-Scale Event-Centric Summarization Dataset for Chinese Multi-News Documents
  20. How do Transformers Learn Implicit Reasoning?
  21. Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
  22. LLMAEL: Large Language Models are Good Context Augmenters for Entity Linking
  23. LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder
  24. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
  25. LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-Context QA
  26. LongReward: Improving Long-context Large Language Models with AI Feedback
  27. LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language Models
  28. LongWriter: Unleashing 10, 000+ Word Generation from Long Context LLMs
  29. Pre-training Distillation for Large Language Models: A Design Space Exploration
  30. RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
  31. SeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval Augmented Generation
  32. Simulating Classroom Education with LLM-Empowered Agents
  33. SoAy: A Solution-based LLM API-using Methodology for Academic Information Seeking
  34. StoryWriter: A Multi-Agent Framework for Long Story Generation
  35. Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
  36. VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
  37. VocQuiz: Vocabulary Question Generation for English Language Education
  38. ADELIE: Aligning Large Language Models on Information Extraction
  39. DiaKoP: Dialogue-based Knowledge-oriented Programming for Neural-symbolic Knowledge Base Question Answering
  40. DocEE-zh: A Fine-grained Benchmark for Chinese Document-level Event Extraction
  41. Event GDR: Event-Centric Generative Document Retrieval
  42. How Proficient Are Large Language Models in Formal Languages? An In-Depth Insight for Knowledge Base Question Answering
  43. KB-Plugin: A Plug-and-play Framework for Large Language Models to Induce Programs over Low-resourced Knowledge Bases
  44. KoLA: Carefully Benchmarking World Knowledge of Large Language Models
  45. LC4EE: LLMs as Good Corrector for Event Extraction
  46. LongAlign: A Recipe for Long Context Alignment of Large Language Models
  47. LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
  48. MAVEN-ARG: Completing the Puzzle of All-in-One Event Understanding Dataset with Event Argument Annotation
  49. MAVEN-FACT: A Large-scale Event Factuality Detection Dataset
  50. MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification
  51. R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
  52. WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models