PPaperPicks

Junyang Lin

52 papers at tracked venues · 40 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
  2. From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning
  3. HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
  4. Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
  5. PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
  6. ToolRM: Towards Agentic Tool-Use Reward Modeling
  7. Towards Better Correctness and Efficiency in Code Generation
  8. UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
  9. A Probabilistic Inference Scaling Theory for LLM Self-Correction
  10. A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
  11. Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models
  12. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
  13. CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention
  14. CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
  15. CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration
  16. Chain of Execution Supervision Promotes General Reasoning in Large Language Models
  17. CodeArena: Evaluating and Aligning CodeLLMs on Human Preference
  18. Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
  19. DataMan: Data Manager for Pre-training Large Language Models
  20. Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
  21. Disentangling Reasoning Tokens and Boilerplate Tokens For Language Model Fine-tuning
  22. Efficient Long Context Fine-tuning with Chunk Flow
  23. FPE2M2: Approaching Lossless and Efficient Quantization with Native Floating Point
  24. Fine-Tuning Language Models with Collaborative and Semantic Experts
  25. Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
  26. HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning
  27. InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
  28. LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
  29. MARGE: Improving Math Reasoning with Guided Exploration
  30. Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control is Easier than You Think
  31. NOVA-63: Native Omni-lingual Versatile Assessments of 63 Disciplines
  32. OpenHands: An Open Platform for AI Software Developers as Generalist Agents
  33. P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
  34. Parallel Scaling Law for Language Models
  35. PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
  36. ProcessBench: Identifying Process Errors in Mathematical Reasoning
  37. Qwen2.5-xCoder: Multi-Agent Collaboration for Multilingual Code Instruction Tuning
  38. RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing
  39. Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
  40. Rethinking Data Selection at Scale: Random Selection is Almost All You Need
  41. Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inference
  42. START: Self-taught Reasoner with Tools
  43. Self-Steering Optimization: Autonomous Preference Optimization for Large Language Models
  44. Synthesizing Software Engineering Data in a Test-Driven Manner
  45. Teaching Language Models to Reason with Tools
  46. The Lessons of Developing Process Reward Models in Mathematical Reasoning
  47. Turning the Tide: Repository-based Code Reflection
  48. #InsTag: Instruction Tagging for Analyzing Supervised Fine-tuning of Large Language Models
  49. An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
  50. Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?
  51. Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models
  52. Synthesizing Text-to-SQL Data from Weak and Strong LLMs