PPaperPicks

Moontae Lee

29 papers at tracked venues · 18 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Adaptive Retrieval for Reasoning
  2. DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
  3. Efficiently Learning To Reason or Not to Reason: Root-token Policy Optimization for Adaptive Thinking
  4. IRPO: Implicit Policy Regularized Preference Optimization
  5. 3D Denoisers Are Good 2D Teachers: Molecular Pretraining via Denoising and Cross-Modal Distillation
  6. Counterfactual Voting Adjustment for Quality Assessment and Fairer Voting in Online Platforms with Helpfulness Evaluation
  7. Learning to Explore and Select for Coverage-Conditioned Retrieval-Augmented Generation
  8. MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
  9. Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews
  10. One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL
  11. Overlapping Context with Variable-Length Stride Increases Diversity when Training Large Language Model for Code
  12. PANORAMA: A Dataset and Benchmarks Capturing Decision Trails and Rationales in Patent Examination
  13. Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
  14. Shifting from Ranking to Set Selection for Retrieval Augmented Generation
  15. The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
  16. Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
  17. Training-free Detection of AI-generated images via Cropping Robustness
  18. Code Models are Zero-shot Precondition Reasoners
  19. Degeneration-free Policy Optimization: RL Fine-Tuning for Language Models without Degeneration
  20. Learning to Unlearn: Instance-Wise Unlearning for Pre-trained Classifiers
  21. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
  22. Prospector: Improving LLM Agents with Self-Asking and Trajectory Ranking
  23. Semantic Skill Grounding for Embodied Instruction-Following in Cross-Domain Environments
  24. Show, Think, and Tell: Thought-Augmented Fine-Tuning of Large Language Models for Video Captioning
  25. Small Language Models Need Strong Verifiers to Self-Correct Reasoning
  26. Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
  27. When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
  28. YTCommentQA: Video Question Answerability in Instructional Videos
  29. You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments