PPaperPicks

Weinan Zhang

Shanghai Jiao Tong University, John Hopcroft Center for Computer Science, China

90 papers at tracked venues · 69 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A Comprehensive Survey of Process Reward Models: Data Generation, Model Construction, and Usage
  2. A Survey of Large Language Model-Based Search Agents
  3. ACE-Router: Generalizing History-Aware Routing from MCP Tools to the Agent Web
  4. Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning
  5. Attribution-Based Analysis and Optimization of Modular Agentic Workflows
  6. ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
  7. ColorBrowserAgent: Complex Long-Horizon Browser Agent with Adaptive Knowledge Evolution
  8. CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
  9. LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
  10. Modular Representation Compression: Adapting LLM Representations for Efficient and Effective Recommendation
  11. Offline Fictitious Self-Play for Competitive Games
  12. PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork
  13. Progra: Progress-Aware Reinforcement Learning for Multi-Turn Function Calling
  14. Sell It Before You Make It: Revolutionizing E-Commerce with Personalized AI-Generated Items
  15. ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
  16. Action First: Leveraging Preference-Aware Actions for More Effective Decision-Making in Interactive Recommender Systems
  17. AdvKT: An Adversarial Multi-step Training Framework for Knowledge Tracing
  18. AgentIR: 2nd Workshop on Agent-based Information Retrieval
  19. AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent Systems
  20. An Automatic Graph Construction Framework based on Large Language Models for Recommendation
  21. Autonomous Goal Detection and Cessation in Reinforcement Learning: A Case Study on Source Term Estimation
  22. Beyond Graph Convolution: Multimodal Recommendation with Topology-aware MLPs
  23. Boost, Disentangle, and Customize: A Robust System2-to-System1 Pipeline for Code Generation
  24. Bursting Filter Bubble: Enhancing Serendipity Recommendations with Aligned Large Language Models
  25. CodePRM: Execution Feedback-enhanced Process Reward Model for Code Generation
  26. ContraDiff: Planning Towards High Return States via Contrastive Learning
  27. D2K: Turning Historical Data into Retrievable Knowledge for Recommender Systems
  28. DebateCoder: Towards Collective Intelligence of LLMs via Test Case Driven LLM Debate for Code Generation
  29. Diffusion Models for Recommender Systems: From Content Distribution To Content Creation
  30. Diversity-Aware Self-Paced Data Selection for LLM Fine-Tuning
  31. DriveGen: Towards Infinite Diverse Traffic Scenarios with Large Models
  32. Efficiency Unleashed: Inference Acceleration for LLM-based Recommender Systems with Speculative Decoding
  33. Fast Second-Order Online Kernel Learning Through Incremental Matrix Sketching and Decomposition
  34. Flexible Realignment of Language Models
  35. GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning
  36. HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Assistant Scenarios
  37. Humanoid Whole-Body Locomotion on Narrow Terrain via Dynamic Balance and Reinforcement Learning
  38. Information-Theoretic Reward Decomposition for Generalizable RLHF
  39. KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills
  40. LLM4CD: Leveraging Large Language Models for Open-World Knowledge Augmented Cognitive Diagnosis
  41. Large Language Models are Demonstration Pre-Selectors for Themselves
  42. Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration
  43. LoopSR: Looping Sim-and-Real for Lifelong Policy Adaptation of Legged Robots
  44. MobileUse: A Hierarchical Reflection-Driven GUI Agent for Autonomous Mobile Operation
  45. NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code Debugging
  46. ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement Learning
  47. Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency
  48. RethinkMCTS: Refining Erroneous Thoughts in Monte Carlo Tree Search for Code Generation
  49. Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
  50. Robust Function-Calling for On-Device Language Model via Function Masking
  51. Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport
  52. Stop DDoS Attacking the Research Community with AI-Generated Survey Papers
  53. ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
  54. Uni-RL: Unifying Online and Offline RL via Implicit Value Regularization
  55. Unleashing the Potential of Multi-Channel Fusion in Retrieval for Personalized Recommendations
  56. Unlocking the Potential of Decentralized LLM-based MAS: Privacy Preservation and Monetization in Collective Intelligence
  57. Why Not Together? A Multiple-Round Recommender System for Queries and Items
  58. World Model-Based Perception for Visual Legged Locomotion
  59. 4DBInfer: A 4D Benchmarking Toolbox for Graph-Centric Predictive Modeling on RDBs
  60. 4DBInfer: A 4D Benchmarking Toolbox for Graph-Centric Predictive Modeling on RDBs
  61. AgentIR: 1st Workshop on Agent-based Information Retrieval
  62. AlignRec: Aligning and Training in Multimodal Recommendations
  63. AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and Training
  64. Boosting Studies of Multi-Agent Reinforcement Learning on Google Research Football Environment: The Past, Present, and Future
  65. CityFlowER: An Efficient and Realistic Traffic Simulator with Embedded Machine Learning Models
  66. ClickPrompt: CTR Models are Strong Prompt Generators for Adapting Language Models to CTR Prediction
  67. DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching
  68. Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning
  69. Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization
  70. DisCo: Towards Harmonious Disentanglement and Collaboration between Tabular and Semantic Space for Recommendation
  71. ELCoRec: Enhance Language Understanding with Co-Propagation of Numerical and Categorical Features for Recommendation
  72. ELF-Gym: Evaluating Large Language Models Generated Features for Tabular Prediction
  73. FLIP: Fine-grained Alignment between ID-based Models and Pretrained Language Models for CTR Prediction
  74. GFS: Graph-based Feature Synthesis for Prediction over Relational Databases
  75. HiFI: Hierarchical Fairness-aware Integrated Ranking with Constrained Reinforcement Learning
  76. InfoRank: Unbiased Learning-to-Rank via Conditional Mutual Information Minimization
  77. K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization
  78. Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
  79. M-scan: A Multi-Scenario Causal-driven Adaptive Network for Recommendation
  80. MADiff: Offline Multi-agent Learning with Diffusion Models
  81. MemoCRS: Memory-enhanced Sequential Conversational Recommender Systems with Large Language Models
  82. ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient Update
  83. ReLLa: Retrieval-enhanced Large Language Models for Lifelong Sequential Behavior Comprehension in Recommendation
  84. Recall-Augmented Ranking: Enhancing Click-Through Rate Prediction Accuracy with Cross-Stage Data
  85. Reinforcing LLM Agents via Policy Optimization with Action Decomposition
  86. SINKT: A Structure-Aware Inductive Knowledge Tracing Model with Large Language Model
  87. TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned Decision
  88. Towards Open-World Recommendation with Knowledge Augmentation from Large Language Models
  89. Vision-Language Foundation Models as Effective Robot Imitators
  90. ZSC-Eval: An Evaluation Toolkit and Benchmark for Multi-agent Zero-shot Coordination