P
PaperPicks
Conferences
← All conferences
Editions
2026
2025
2024
MLSys
2024
N
Conference on Machine Learning and Systems
Official site ↗
39
accepted papers
39
authors
May 12-16, 2024
dates
USA
location
39
/ 39 papers
All
39
Main Track
37
Editorship
2
1
Accelerating ReLU for MPC-Based Private Inference with a Communication-Efficient Sign Estimation
Kiwan Maeng
DBLP profile ↗
ORCID search ↗
2
Accurate Low-Degree Polynomial Approximation of Non-Polynomial Operators for Fast Private Inference in Homomorphic Encryption
Jingtian Dang
DBLP profile ↗
ORCID search ↗
3
ACROBAT: Optimizing Auto-batching of Dynamic Deep Learning at Compile Time
Pratik Fegade
DBLP profile ↗
ORCID search ↗
4
Atom: Low-Bit Quantization for Efficient and Accurate LLM Serving
Yilong Zhao
DBLP profile ↗
ORCID search ↗
5
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
Ji Lin
DBLP profile ↗
ORCID search ↗
6
CloudEval-YAML: A Practical Benchmark for Cloud Configuration Generation
Yifei Xu
DBLP profile ↗
ORCID search ↗
7
COMET: Neural Cost Model Explanation Framework
Isha Chaudhary
DBLP profile ↗
ORCID search ↗
8
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
Ye Tian
DBLP profile ↗
ORCID search ↗
9
Disaggregated Multi-Tower: Topology-aware Modeling Technique for Efficient Large Scale Recommendation
Liang Luo
DBLP profile ↗
ORCID search ↗
10
Distributed Matrix-Based Sampling for Graph Neural Network Training
Alok Tripathy
DBLP profile ↗
ORCID search ↗
11
Does Compressing Activations Help Model Parallel Training?
Song Bian
DBLP profile ↗
ORCID search ↗
12
Efficient Post-training Quantization with FP8 Formats
Haihao Shen
DBLP profile ↗
ORCID search ↗
13
FedTrans: Efficient Federated Learning via Multi-Model Transformation
Yuxuan Zhu
DBLP profile ↗
ORCID search ↗
14
Fine-Tuning Language Models Using Formal Methods Feedback: A Use Case in Autonomous Systems
Yunhao Yang
DBLP profile ↗
ORCID search ↗
15
FLASH: Fast Model Adaptation in ML-Centric Cloud Platforms
Haoran Qiu
DBLP profile ↗
ORCID search ↗
16
FlashDecoding++: Faster Large Language Model Inference with Asynchronization, Flat GEMM Optimization, and Heuristics
Ke Hong
DBLP profile ↗
ORCID search ↗
17
HeteGen: Efficient Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
Xuanlei Zhao
DBLP profile ↗
ORCID search ↗
18
HeteroSwitch: Characterizing and Taming System-Induced Data Heterogeneity in Federated Learning
Gyudong Kim
DBLP profile ↗
ORCID search ↗
19
JIT-Q: Just-in-time Quantization with Processing-In-Memory for Efficient ML Training
Mohamed Assem Ibrahim
DBLP profile ↗
ORCID search ↗
20
Keyformer: KV Cache reduction through key tokens selection for Efficient Generative Inference
Muhammad Adnan
DBLP profile ↗
ORCID search ↗
21
L-GreCo: Layerwise-adaptive Gradient Compression For Efficient Data-parallel Deep Learning
Ilia Markov
DBLP profile ↗
ORCID search ↗
22
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
Chenyu Jiang
DBLP profile ↗
ORCID search ↗
23
LIFL: A Lightweight, Event-driven Serverless Platform for Federated Learning
Shixiong Qi
DBLP profile ↗
ORCID search ↗
24
On Latency Predictors for Neural Architecture Search
Yash Akhauri
DBLP profile ↗
ORCID search ↗
25
Proceedings of the Eighth Conference on Machine Learning and Systems, MLSys 2025, Santa Clara, CA, USA, May 12-15, 2025
Matei Zaharia
DBLP profile ↗
ORCID 0000-0002-7547-7204 ↗
26
Proceedings of the Seventh Annual Conference on Machine Learning and Systems, MLSys 2024, Santa Clara, CA, USA, May 13-16, 2024
Phillip B. Gibbons
DBLP profile ↗
ORCID search ↗
27
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
In Gim
DBLP profile ↗
ORCID search ↗
28
Proteus: Preserving Model Confidentiality during Graph Optimizations
Yubo Gao
DBLP profile ↗
ORCID search ↗
29
Punica: Multi-Tenant LoRA Serving
Lequn Chen
DBLP profile ↗
ORCID search ↗
30
Q-Hitter: A Better Token Oracle for Efficient LLM Inference via Sparse-Quantized KV Cache
Zhenyu Zhang
DBLP profile ↗
ORCID search ↗
31
QMoE: Sub-1-Bit Compression of Trillion Parameter Models
Elias Frantar
DBLP profile ↗
ORCID search ↗
32
Schrodinger's FP Training Neural Networks with Dynamic Floating-Point Containers
Milos Nikolic
DBLP profile ↗
ORCID search ↗
33
SiDA: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
Zhixu Du
DBLP profile ↗
ORCID search ↗
34
SLoRA: Scalable Serving of Thousands of LoRA Adapters
Ying Sheng
DBLP profile ↗
ORCID search ↗
35
Torch2Chip: An End-to-end Customizable Deep Neural Network Compression and Deployment Toolkit for Prototype Hardware Accelerator Design
Jian Meng
DBLP profile ↗
ORCID search ↗
36
UniDM: A Unified Framework for Data Manipulation with Large Language Models
Yichen Qian
DBLP profile ↗
ORCID search ↗
37
VIDUR: A Large-Scale Simulation Framework for LLM Inference
Amey Agrawal
DBLP profile ↗
ORCID search ↗
38
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
Size Zheng
DBLP profile ↗
ORCID search ↗
39
VQPy: An Object-Oriented Approach to Modern Video Analytics
Shan Yu
DBLP profile ↗
ORCID search ↗