PPaperPicks

Xipeng Qiu

84 papers at tracked venues · 56 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
  2. ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
  3. Efficient KL Divergence Estimation via Truncated Top-K Integration for Large Language Models
  4. HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
  5. How to Set the Learning Rate for Large-Scale Pre-training?
  6. LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
  7. Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
  8. VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
  9. WESR: A Benchmark and Strong Baseline for Word-level Event-Speech Recognition
  10. XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
  11. AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments
  12. Are LLMs Rational Investors? A Study on the Financial Bias in LLMs
  13. BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments
  14. CAMIEval: Enhancing NLG Evaluation through Multidimensional Comparative Instruction-Following Analysis
  15. CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
  16. ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement Learning
  17. CritiQ: Mining Data Quality Criteria from Human Preferences
  18. Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
  19. Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLMs
  20. Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
  21. Dynamic and Generalizable Process Reward Modeling
  22. Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
  23. FastMCTS: A Simple Sampling Strategy for Data Synthesis
  24. FiNE: Filtering and Improving Noisy Data Elaborately with Large Language Models
  25. Firewall Routing: Blocking Leads to Better Hybrid Inference for LLMs
  26. ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation Detection
  27. How to Mitigate Overfitting in Weak-to-strong Generalization?
  28. INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
  29. Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
  30. MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
  31. MetaAlign: Align Large Language Models with Diverse Preferences during Inference Time
  32. Multi-Programming Language Sandbox for LLMs
  33. Pre-Trained Policy Discriminators are General Reward Models
  34. Prior-Fitted Networks Scale to Larger Datasets When Treated as Weak Learners
  35. ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning
  36. R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
  37. REARANK: Reasoning Re-ranking Agent via Reinforcement Learning
  38. ReAttention: Training-Free Infinite Context with Finite Attention Scope
  39. Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
  40. Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Models
  41. Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
  42. Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
  43. UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
  44. UnitCoder: Scalable Code Synthesis from Pre-training Corpora
  45. VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
  46. VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
  47. VideoRoPE: What Makes for Good Video Rotary Position Embedding?
  48. VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
  49. World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
  50. World-aware Planning Narratives Enhance Large Vision-Language Model Planner
  51. AdaLomo: Low-memory Optimization with Adaptive Learning Rate
  52. Alignment for Honesty
  53. AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
  54. Balanced Data Sampling for Language Model Training with Clustering
  55. Calibrating the Confidence of Large Language Models by Eliciting Fidelity
  56. Can AI Assistants Know What They Don't Know?
  57. Can Language Models Learn to Skip Steps?
  58. Code Needs Comments: Enhancing Code LLMs with Comment Augmentation
  59. DenoSent: A Denoising Objective for Self-Supervised Sentence Representation Learning
  60. Enhancing EEG-to-Text Decoding through Transferable Representations from Pre-trained Contrastive EEG-Text Masked Autoencoder
  61. Explicit Memory Learning with Expectation Maximization
  62. F-Eval: Asssessing Fundamental Abilities with Refined Evaluation Methods
  63. Flames: Benchmarking Value Alignment of LLMs in Chinese
  64. Full Parameter Fine-tuning for Large Language Models with Limited Resources
  65. GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
  66. Identifying Semantic Induction Heads to Understand In-Context Learning
  67. InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
  68. Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation
  69. L-Eval: Instituting Standardized Evaluation for Long Context Language Models
  70. LLM can Achieve Self-Regulation via Hyperparameter Aware Generation
  71. LLatrieval: LLM-Verified Retrieval for Verifiable Generation
  72. LongWanjuan: Towards Systematic Measurement for Long Text Quality
  73. Making Large Language Models Better Reasoners with Orchestrated Streaming Experiences
  74. Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental Chunk
  75. Pixel-Level Semantic Correspondence Through Layout-Aware Representation Learning and Multi-Scale Matching Integration
  76. Reasoning in Flux: Enhancing Large Language Models Reasoning through Uncertainty-aware Adaptive Guidance
  77. R³-NL2GQL: A Model Coordination and Knowledge Graph Alignment Approach for NL2GQL
  78. Scaling Laws for Fact Memorization of Large Language Models
  79. Scaling Laws of RoPE-based Extrapolation
  80. SpeechAlign: Aligning Speech Generation to Human Preferences
  81. SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models
  82. Training-Free Long-Context Scaling of Large Language Models
  83. Turn Waste into Worth: Rectifying Top-k Router of MoE
  84. Unified Active Retrieval for Retrieval Augmented Generation