P
PaperPicks
Conferences
Zhenyu Zhang
14 papers at tracked venues · 12 at CORE A* · active 2024–2025
DBLP profile ↗
ORCID search ↗
Venues
ICML
×7
ICLR
×3
MLSys
×2
AAAI
×1
NeurIPS
×1
Frequent coauthors
Hanqing Zhu
DBLP profile ↗
ORCID search ↗
×1
Xialie Zhuang
DBLP profile ↗
ORCID search ↗
×1
Yeonju Ro
DBLP profile ↗
ORCID search ↗
×1
Yuxin Zhang
DBLP profile ↗
ORCID search ↗
×1
Jiawei Zhao
DBLP profile ↗
ORCID search ↗
×1
Harry Dong
DBLP profile ↗
ORCID search ↗
×1
Yuandong Tian
DBLP profile ↗
ORCID search ↗
×1
Pingzhi Li
DBLP profile ↗
ORCID search ↗
×1
Lu Yin
DBLP profile ↗
ORCID search ↗
×1
Zhangheng Li
DBLP profile ↗
ORCID search ↗
×1
Zhen Tan
DBLP profile ↗
ORCID search ↗
×1
Papers
APOLLO: SGD-like Memory, AdamW-level Performance
MLSys 2025
·
Hanqing Zhu
DBLP profile ↗
ORCID search ↗
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
ICML 2025
·
Xialie Zhuang
DBLP profile ↗
ORCID search ↗
On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention for Long-Context LLM Serving
ICML 2025
·
Yeonju Ro
DBLP profile ↗
ORCID search ↗
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
ICLR 2025
·
Zhenyu Zhang
CaM: Cache Merging for Memory-efficient LLMs Inference
ICML 2024
·
Yuxin Zhang
DBLP profile ↗
ORCID search ↗
Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional Encoding
NeurIPS 2024
·
Zhenyu Zhang
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
ICML 2024
·
Jiawei Zhao
DBLP profile ↗
ORCID search ↗
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
ICML 2024
·
Harry Dong
DBLP profile ↗
ORCID search ↗
JoMA: Demystifying Multilayer Transformers via Joint Dynamics of MLP and Attention
ICLR 2024
·
Yuandong Tian
DBLP profile ↗
ORCID search ↗
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
ICLR 2024
·
Pingzhi Li
DBLP profile ↗
ORCID search ↗
Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
ICML 2024
·
Lu Yin
DBLP profile ↗
ORCID search ↗
Q-Hitter: A Better Token Oracle for Efficient LLM Inference via Sparse-Quantized KV Cache
MLSys 2024
·
Zhenyu Zhang
Sparse Cocktail: Every Sparse Pattern Every Sparse Ratio All At Once
ICML 2024
·
Zhangheng Li
DBLP profile ↗
ORCID search ↗
Sparsity-Guided Holistic Explanation for LLMs with Interpretable Inference-Time Intervention
AAAI 2024
·
Zhen Tan
DBLP profile ↗
ORCID search ↗