PPaperPicks

Fei Huang

Alibaba Group, DAMO Academy, Language Technologies Lab, Hangzhou, China

119 papers at tracked venues · 80 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
  2. Efficient and Effective In-context Demonstration Selection with Coreset
  3. EvoRoute: Experience-Driven Self-Routing LLM Agent Systems
  4. Experience-driven Multi-turn Reinforcement Learning for GUI Agents
  5. FishFlow: A LLM-Empowered Dynamic Pricing Framework for Online Fleamarket Platform
  6. Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format
  7. MOA: Multi-Objective Alignment for Role-Playing Agents
  8. MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
  9. Memp: Exploring Agent Procedural Memory
  10. Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
  11. RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
  12. Reasoning-Guided Exploration for Online DPO
  13. STORM: A Spatio-Temporal Factor Model Based on Dual Vector Quantized Variational Autoencoders for Financial Trading
  14. Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration
  15. Selective Weak-to-Strong Generalization
  16. Towards General Agentic Intelligence via Environment Scaling
  17. Understanding Generalization in Role-Playing Models via Information Theory
  18. Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
  19. AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
  20. Agentic Knowledgeable Self-awareness
  21. Benchmarking Agentic Workflow Generation
  22. Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
  23. CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention
  24. CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization
  25. ConText: Driving In-context Learning for Text Removal and Segmentation
  26. CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis
  27. DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
  28. DISC: Plug-and-Play Decoding Intervention with Similarity of Characters for Chinese Spelling Check
  29. Debate Helps Weak-to-Strong Generalization
  30. DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modeling
  31. DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking
  32. Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference
  33. EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
  34. EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning
  35. Effective Two-Stage Knowledge Transfer for Multi-Entity Cross-Domain Recommendation
  36. Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training
  37. EvolveSearch: An Iterative Self-Evolving Search Agent
  38. Exploiting Presentative Feature Distributions for Parameter-Efficient Continual Learning of Large Language Models
  39. ExploraCoder: Advancing Code Generation for Multiple Unseen APIs via Planning and Chained Exploration
  40. FishBargain: An LLM-Empowered Bargaining Agent for Online Fleamarket Platform Sellers
  41. GSID: Generative Semantic Indexing for E-Commerce Product Understanding
  42. IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
  43. IU4Rec: Interest Unit-Based Product Organization and Recommendation for E-Commerce Platform
  44. KBM: Delineating Knowledge Boundary for Adaptive Retrieval in Large Language Models
  45. LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs - No Silver Bullet for LC or RAG Routing
  46. LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
  47. Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
  48. MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
  49. Multi-Value-Product Retrieval-Augmented Generation for Industrial Product Attribute Value Identification
  50. OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
  51. OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking
  52. On the Role of Attention Heads in Large Language Model Safety
  53. OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis
  54. P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
  55. PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
  56. ProductAgent: Benchmarking Conversational Product Search Agent with Asking Clarification Questions
  57. RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing
  58. Reverse Preference Optimization for Complex Instruction Following
  59. SDPO: Segment-Level Direct Preference Optimization for Social Agents
  60. Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
  61. StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization
  62. Supervised Optimism Correction: Be Confident When LLMs Are Sure
  63. Supportiveness-based Knowledge Rewriting for Retrieval-augmented Language Modeling
  64. SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization
  65. SynWorld: Virtual Scenario Synthesis for Agentic Action Knowledge Refinement
  66. Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
  67. Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
  68. Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization
  69. VLM-R³: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
  70. VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
  71. WebDancer: Towards Autonomous Information Seeking Agency
  72. WebWalker: Benchmarking LLMs in Web Traversal
  73. WritingBench: A Comprehensive Benchmark for Generative Writing
  74. mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
  75. mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
  76. A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language Models
  77. Agent Planning with World Knowledge Model
  78. Browse and Concentrate: Comprehending Multimodal Content via Prior-LLM Context Fusion
  79. Budget-Constrained Tool Learning with Planning
  80. DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
  81. EcomGPT: Instruction-Tuning Large Language Models with Chain-of-Task Tasks for E-commerce
  82. EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
  83. Exploring Key Point Analysis with Pairwise Generation and Graph Partitioning
  84. FactCHD: Benchmarking Fact-Conflicting Hallucination Detection
  85. Fine-Tuning Language Models with Reward Learning on Policy
  86. FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
  87. Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use
  88. Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
  89. How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
  90. IPL: Leveraging Multimodal Large Language Models for Intelligent Product Listing
  91. Improving Factual Consistency of News Summarization by Contrastive Preference Optimization
  92. Improving Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts
  93. Iterative Forward Tuning Boosts In-Context Learning in Language Models
  94. Knowledge Mechanisms in Large Language Models: A Survey and Perspective
  95. Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
  96. Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
  97. MIBench: Evaluating Multimodal Large Language Models over Multiple Images
  98. MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
  99. Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
  100. Model Composition for Multimodal Large Language Models
  101. OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
  102. One-Shot Learning as Instruction Data Prospector for Large Language Models
  103. PANDA: Preference Adaptation for Enhancing Domain-Specific Abilities of LLMs
  104. Platypus: A Generalized Specialist Model for Reading Text in Various Forms
  105. Preference Ranking Optimization for Human Alignment
  106. Query Routing for Homogeneous Tools: An Instantiation in the RAG Scenario
  107. RaFe: Ranking Feedback Improves Query Rewriting for RAG
  108. Retrieved In-Context Principles from Previous Mistakes
  109. Self-Retrieval: End-to-End Information Retrieval with One Large Language Model
  110. SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence Understanding
  111. Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
  112. SocialBench: Sociality Evaluation of Role-Playing Conversational Agents
  113. TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token Merging
  114. Visual Text Generation in the Wild
  115. WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models
  116. mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval
  117. mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
  118. mPLUG-OwI2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
  119. mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model