PPaperPicks

Dan Alistarh

IST Austria, Klosterneuburg, Austria

33 papers at tracked venues · 25 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Speculative Decoding Speed-of-Light: Optimal Lower Bounds via Branching Random Walks
  2. "Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
  3. Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
  4. Efficient Data Selection at Scale via Influence Distillation
  5. EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
  6. HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
  7. HIGGS: Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
  8. Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
  9. Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
  10. LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
  11. Layer-wise Quantization for Quantized Optimistic Dual Averaging
  12. QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
  13. Quartet: Native FP4 Training Can Be Optimal for Large Language Models
  14. Scalable Mechanistic Neural Networks
  15. The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
  16. Unified Scaling Laws for Compressed Representations
  17. Wasserstein Distances, Neuronal Entanglement, and Sparsity
  18. AsGrad: A Sharp Unified Analysis of Asynchronous-SGD Algorithms
  19. Communication-Efficient Federated Learning With Data and Client Heterogeneity
  20. Error Feedback Can Accurately Compress Preconditioners
  21. Extreme Compression of Large Language Models via Additive Quantization
  22. L-GreCo: Layerwise-adaptive Gradient Compression For Efficient Data-parallel Deep Learning
  23. Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
  24. MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence
  25. PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
  26. QMoE: Sub-1-Bit Compression of Trillion Parameter Models
  27. QUIK: Towards End-to-end 4-Bit Inference on Generative Large Language Models
  28. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
  29. RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
  30. SPADE: Sparsity-Guided Debugging for Deep Neural Networks
  31. Scaling Laws for Sparsely-Connected Foundation Models
  32. SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
  33. The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information