PPaperPicks

Maosong Sun

Tsinghua University, Department of Computer Science and Technology, Beijing, China

116 papers at tracked venues · 91 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Task
  2. AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
  3. AutoVecCoder: Teaching LLMs to Generate Explicitly Vectorized Code
  4. CheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented Reasoning
  5. Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
  6. EchoMLLM: Incentivizing Echocardiographic Video Understanding with Keyframe Grounding and Report Generation
  7. Empirical Analysis of Decoding Biases in Masked Diffusion Models
  8. Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection
  9. FaithLens: Detecting and Explaining Faithfulness Hallucination
  10. From Scaffolding to Assimilation: Progressive Structural Internalization for Format-Constrained Creative Text Generation
  11. IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
  12. ImCoref-CeS: An Improved Lightweight Pipeline for Coreference Resolution with LLM-based Checker-Splitter Refinement
  13. LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic Pyramid
  14. Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
  15. MEIC-DT: Memory-Efficient Incremental Clustering for Long-Text Coreference Resolution with Dual-Threshold Constraints
  16. MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization
  17. Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation
  18. Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
  19. RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation
  20. Revealing the Attention Floating Mechanism in Masked Diffusion Models
  21. StateX: Enhancing RNN Recall via Post-training State Expansion
  22. Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning
  23. UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
  24. A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
  25. A Top-down Graph-based Tool for Modeling Classical Semantic Maps: A Case Study of Supplementary Adverbs
  26. A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings
  27. APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
  28. AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
  29. Advancing LLM Reasoning Generalists with Preference Trees
  30. AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
  31. AgentRM: Enhancing Agent Generalization with Reward Modeling
  32. Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
  33. AutoClean: LLMs Can Prepare Their Training Corpus
  34. Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
  35. CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
  36. CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
  37. ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
  38. ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation
  39. Cost-Optimal Grouped-Query Attention for Long-Context Modeling
  40. DCAD-2000: A Multilingual Dataset across 2000+ Languages with Data Cleaning as Anomaly Detection
  41. DeepNote: Note-Centric Deep Retrieval-Augmented Generation
  42. Document Segmentation Matters for Retrieval-Augmented Generation
  43. Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub
  44. FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
  45. From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
  46. Fusing Highly Specialized Language Models for Comprehensive Expertise
  47. GATEAU: Selecting Influential Samples for Long Context Alignment
  48. GLTW: Joint Improved Graph Transformer and LLM via Three-Word Language for Knowledge Graph Completion
  49. GUICourse: From General Vision Language Model to Versatile GUI Agent
  50. Internet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative Intelligence
  51. Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
  52. KBAlign: Efficient Self Adaptation on Specific Textual Knowledge Bases
  53. LLM×MapReduce-V3: Enabling Interactive In-Depth Survey Generation through a MCP-Driven Hierarchically Modular Agent System
  54. LLM×MapReduce: Simplified Long-Sequence Processing using Large Language Models
  55. Learning to Generate Structured Output with Schema Reinforcement Learning
  56. Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
  57. Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
  58. Multi-Agent Collaboration via Evolving Orchestration
  59. NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms
  60. On LLM-Based Scientific Inductive Reasoning Beyond Equations
  61. Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System
  62. ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
  63. PersLLM: A Personified Training Approach for Large Language Models
  64. Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance
  65. RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
  66. RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework
  67. RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
  68. Rational Decision-Making Agent with Learning Internal Utility Judgment
  69. Scaling Large Language Model-based Multi-Agent Collaboration
  70. Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model Training
  71. Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
  72. The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE Training
  73. The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
  74. TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
  75. Value Compass Benchmarks: A Comprehensive, Generative and Self-Evolving Platform for LLMs' Value Evaluation
  76. VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
  77. WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models
  78. XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
  79. AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
  80. Beyond Natural Language: LLMs Leveraging Alternative Formats for Enhanced Reasoning and Communication
  81. Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
  82. Browse and Concentrate: Comprehending Multimodal Content via Prior-LLM Context Fusion
  83. CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language Models
  84. Can Large Language Models Analyze Graphs like Professionals? A Benchmark, Datasets and Models
  85. ChatDev: Communicative Agents for Software Development
  86. Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
  87. DebugBench: Evaluating Debugging Capability of Large Language Models
  88. DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models
  89. Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
  90. Empowering Private Tutoring by Chaining Large Language Models
  91. Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate Pairs
  92. Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages
  93. Experiential Co-Learning of Software-Developing Agents
  94. Exploring the Benefit of Activation Sparsity in Pre-training
  95. Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics
  96. FastFiD: Improve Inference Efficiency of Open Domain Question Answering via Sentence Selection
  97. InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
  98. LEGENT: Open Platform for Embodied Agents
  99. Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
  100. LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks
  101. MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data Visualization
  102. Model Composition for Multimodal Large Language Models
  103. OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
  104. On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
  105. Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
  106. Predicting Emergent Abilities with Infinite Resolution Evaluation
  107. RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-Grained Correctional Human Feedback
  108. Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models
  109. RepoAgent: An LLM-Powered Open-Source Framework for Repository-level Code Documentation Generation
  110. StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models
  111. Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
  112. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
  113. ULTRAFEEDBACK: Boosting Language Models with Scaled AI Feedback
  114. UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs
  115. UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset
  116. ınftyBench: Extending Long Context Evaluation Beyond 100K Tokens