P
PaperPicks
Conferences
← All conferences
Editions
2026
2025
2024
MLSys
2025
N
Conference on Machine Learning and Systems
Official site ↗
61
accepted papers
61
authors
May 13-16, 2025
dates
USA
location
61
/ 61 papers
All
61
Main Track
61
1
A Bring-Your-Own-Model Approach for ML-Driven Storage Placement in Warehouse-Scale Computers
Chenxi Yang
DBLP profile ↗
ORCID search ↗
2
AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine
Carlo Siebenschuh
DBLP profile ↗
ORCID search ↗
3
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
Zhiqiang Xie
DBLP profile ↗
ORCID search ↗
4
AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
Yinfang Chen
DBLP profile ↗
ORCID search ↗
5
APOLLO: SGD-like Memory, AdamW-level Performance
Hanqing Zhu
DBLP profile ↗
ORCID search ↗
6
Balancing Pipeline Parallelism with Vocabulary Parallelism
Man Tsung Yeung
DBLP profile ↗
ORCID search ↗
7
COMET: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
Shulai Zhang
DBLP profile ↗
ORCID search ↗
8
Context Parallelism for Scalable Million-Token Inference
Amy Yang
DBLP profile ↗
ORCID search ↗
9
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
Sohaib Ahmad
DBLP profile ↗
ORCID search ↗
10
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
Marco Federici
DBLP profile ↗
ORCID search ↗
11
Efficient On-Device Machine Learning with a Biologically-Plausible Forward-Only Algorithm
Baichuan Huang
DBLP profile ↗
ORCID search ↗
12
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
Geonhwa Jeong
DBLP profile ↗
ORCID search ↗
13
FastTree: Optimizing Attention Kernel and Runtime for Tree-Structured LLM Inference
Zaifeng Pan
DBLP profile ↗
ORCID search ↗
14
FedProphet: Memory-Efficient Federated Adversarial Training via Robust and Consistent Cascade Learning
Minxue Tang
DBLP profile ↗
ORCID search ↗
15
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
Zihao Ye
DBLP profile ↗
ORCID search ↗
16
FlexAttention: A Programming Model for Generating Fused Attention Variants
Juechu Dong
DBLP profile ↗
ORCID search ↗
17
FlexInfer: Flexible LLM Inference with CPU Computations
Seonjin Na
DBLP profile ↗
ORCID search ↗
18
FLStore: Efficient Federated Learning Storage for non-training workloads
Ahmad Faraz Khan
DBLP profile ↗
ORCID search ↗
19
Graph Learning at Scale: Characterizing and Optimizing Pre-Propagation GNNs
Zichao Yue
DBLP profile ↗
ORCID search ↗
20
HyC-LoRA: Memory Efficient LoRA Fine-tuning with Hybrid Activation Compression
Yujin Wang
DBLP profile ↗
ORCID search ↗
21
Interference-aware Edge Runtime Prediction with Conformal Matrix Completion
Tianshu Huang
DBLP profile ↗
ORCID search ↗
22
Know Where You're Uncertain When Planning with Multimodal Foundation Models: A Formal Framework
Neel P. Bhatt
DBLP profile ↗
ORCID search ↗
23
LAVA: Lifetime-Aware VM Allocation with Learned Distributions and Adaptation to Mispredictions
Jianheng Ling
DBLP profile ↗
ORCID search ↗
24
LeanAttention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
Rya Sanovar
DBLP profile ↗
ORCID search ↗
25
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
Francesco Daghero
DBLP profile ↗
ORCID search ↗
26
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
Shang Yang
DBLP profile ↗
ORCID search ↗
27
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
Mingyu Liang
DBLP profile ↗
ORCID search ↗
28
Marconi: Prefix Caching for the Era of Hybrid LLMs
Rui Pan
DBLP profile ↗
ORCID search ↗
29
Mas-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-constrained Edge Devices
Mohammadali Shakerdargah
DBLP profile ↗
ORCID search ↗
30
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
Abhishek Moitra
DBLP profile ↗
ORCID search ↗
31
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
Beichen Huang
DBLP profile ↗
ORCID search ↗
32
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
Xuanlin Jiang
DBLP profile ↗
ORCID search ↗
33
On Distributed Larger-Than-Memory Subset Selection With Pairwise Submodular Functions
Maximilian Böther
DBLP profile ↗
ORCID search ↗
34
Optimizing LLM Queries in Relational Data Analytics Workloads
Shu Liu
DBLP profile ↗
ORCID search ↗
35
Photon: Federated LLM Pre-Training
Lorenzo Sani
DBLP profile ↗
ORCID search ↗
36
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
Daiyaan Arfeen
DBLP profile ↗
ORCID search ↗
37
ProtoRAIL: A Risk-cognizant Imitation Agent for Adaptive vCPU Oversubscription In the Cloud
Lu Wang
DBLP profile ↗
ORCID 0000-0002-7305-1496 ↗
38
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Yujun Lin
DBLP profile ↗
ORCID search ↗
39
Radius: Range-based Gradient Sparsity for Large Foundation Model Pre-training
Mingkai Zheng
DBLP profile ↗
ORCID search ↗
40
ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation
Zhiyu Mei
DBLP profile ↗
ORCID search ↗
41
Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
Wei Gao
DBLP profile ↗
ORCID search ↗
42
Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
Xinyi Zhang
DBLP profile ↗
ORCID search ↗
43
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
Qianchao Zhu
DBLP profile ↗
ORCID search ↗
44
ScaleFusion: Scalable Inference of Spatial-Temporal Diffusion Transformers for High-Resolution Long Video Generation
Jiacheng Yang
DBLP profile ↗
ORCID search ↗
45
Scaling Deep Learning Training with MPMD Pipeline Parallelism
Anxhelo Xhebraj
DBLP profile ↗
ORCID search ↗
46
Seesaw: High-throughput LLM Inference via Model Re-sharding
Qidong Su
DBLP profile ↗
ORCID search ↗
47
Self-Data Distillation for Recovering Quality in Pruned Large Language Models
Vithursan Thangarasa
DBLP profile ↗
ORCID search ↗
48
SOLA: Optimizing SLO Attainment for Large Language Model Serving with State-Aware Scheduling
Ke Hong
DBLP profile ↗
ORCID search ↗
49
Spa: Scaling Graph Neural Network Training on Large graphs via Probabilistic splitting
Sandeep Polisetty
DBLP profile ↗
ORCID search ↗
50
SparseTransX: Efficient Training of Translation-Based Knowledge Graph Embeddings Using Sparse Matrix Operations
Md. Saidul Hoque Anik
DBLP profile ↗
ORCID search ↗
Show 100 more
(11 left)