PPaperPicks

Zhenyu Zhang

14 papers at tracked venues · 12 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. APOLLO: SGD-like Memory, AdamW-level Performance
  2. Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
  3. On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention for Long-Context LLM Serving
  4. R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
    ICLR 2025 · Zhenyu Zhang
  5. CaM: Cache Merging for Memory-efficient LLMs Inference
  6. Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional Encoding
    NeurIPS 2024 · Zhenyu Zhang
  7. GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
  8. Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
  9. JoMA: Demystifying Multilayer Transformers via Joint Dynamics of MLP and Attention
  10. Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
  11. Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
  12. Q-Hitter: A Better Token Oracle for Efficient LLM Inference via Sparse-Quantized KV Cache
    MLSys 2024 · Zhenyu Zhang
  13. Sparse Cocktail: Every Sparse Pattern Every Sparse Ratio All At Once
  14. Sparsity-Guided Holistic Explanation for LLMs with Interpretable Inference-Time Intervention