PPaperPicks

Xunliang Cai

63 papers at tracked venues · 48 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
  2. AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
  3. Attribution-Based Analysis and Optimization of Modular Agentic Workflows
  4. BAPO: Boundary-Aware Policy Optimization for Reliable Agentic Search
  5. CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
  6. Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
  7. DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
  8. Harmonizing Dense and Sparse Signals in Multi-turn RL: Dual-Horizon Credit Assignment for Industrial Sales Agents
  9. How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
  10. LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance
  11. Large-Scale Diverse Synthesis for Mid-Training
  12. LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
  13. MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
  14. MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
  15. Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
  16. PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
  17. Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
  18. Scaling and Transferability of Annealing Strategies in Large Language Model Training
  19. Turning Failures into Value: Negative Experience Replay for RLVR via Confidence Gating and Boundary Failure Sampling
  20. Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
  21. VANE: Guiding High-Value Exploration in RLVR via Outcome-Process Novelty Shaping
  22. WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
  23. A Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy
  24. AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models
  25. AgentRefine: Enhancing Agent Generalization through Refinement Tuning
  26. Don't Half-listen: Capturing Key-part Information in Continual Instruction Tuning
  27. Dynamic Fisher-weighted Model Merging via Bayesian Optimization
  28. Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective
  29. Enhancing LLMs via High-Knowledge Data Selection
  30. FIRE: Flexible Integration of Data Quality Ratings for Effective Pretraining
  31. FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy
  32. Instance-level Randomization: Toward More Stable LLM Evaluations
  33. Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
  34. Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration
  35. Leveraging Unpaired Feedback for Long-Term LLM-based Recommendation Tuning
  36. LogicPro: Improving Complex Logical Reasoning via Program-Guided Learning
  37. MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
  38. Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
  39. Multi-Programming Language Sandbox for LLMs
  40. NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables
  41. Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
  42. PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
  43. Prejudge-Before-Think: Enhancing Large Language Models at Test-Time by Process Prejudge Reasoning
  44. ReMamba: Equip Mamba with Effective Long-Sequence Modeling
  45. Revisit Self-Debugging with Self-Generated Tests for Code Generation
  46. Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies
  47. S3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical Reasoners
  48. SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models
  49. SampleMix: A Sample-wise Pre-training Data Mixing Strategy by Coordinating Data Quality and Diversity
  50. The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
  51. Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs
  52. When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
  53. Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
  54. DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning
  55. Graph-Structured Speculative Decoding
  56. Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
  57. How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data
  58. Learning or Self-aligning? Rethinking Instruction Fine-tuning
  59. Not All Contexts Are Equal: Teaching LLMs Credibility-aware Generation
  60. Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning
  61. Rethinking the Reversal Curse of LLMs: a Prescription from Human Knowledge Reversal
  62. Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
  63. What Makes Quantization for Large Language Model Hard? An Empirical Study from the Lens of Perturbation