PPaperPicks

Yasheng Wang

34 papers at tracked venues · 25 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
  2. ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool learning
  3. Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger
  4. AdvKT: An Adversarial Multi-step Training Framework for Knowledge Tracing
  5. Benchmarking Retrieval-Augmented Multimomal Generation for Document Question Answering
  6. Boost, Disentangle, and Customize: A Robust System2-to-System1 Pipeline for Code Generation
  7. Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
  8. Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step
  9. CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
  10. CodePRM: Execution Feedback-enhanced Process Reward Model for Code Generation
  11. Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
  12. DebateCoder: Towards Collective Intelligence of LLMs via Test Case Driven LLM Debate for Code Generation
  13. DeepDiver: Adaptive Web-Search Intensity Scaling via Reinforcement Learning
  14. Flat-LoRA: Low-Rank Adaptation over a Flat Loss Landscape
  15. Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
  16. Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction
  17. Learning Evolving Tools for Large Language Models
  18. NILE: Internal Consistency Alignment in Large Language Models
  19. NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code Debugging
  20. Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance
  21. QFFT, Question-Free Fine-Tuning for Adaptive Reasoning
  22. RethinkMCTS: Refining Erroneous Thoughts in Monte Carlo Tree Search for Code Generation
  23. RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
  24. RidgeLoRA: Matrix Ridge Enhanced Low-Rank Adaptation of Large Language Models
  25. Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
  26. Spa-Bench: a comprehensive Benchmark for Smartphone Agent Evaluation
  27. Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
  28. ToolACE: Winning the Points of LLM Function Calling
  29. ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
  30. Dynamic Stochastic Decoding Strategy for Open-Domain Dialogue Generation
  31. Evaluating Robustness of Generative Search Engine on Adversarial Factoid Questions
  32. Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios
  33. ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models
  34. SINKT: A Structure-Aware Inductive Knowledge Tracing Model with Large Language Model