PPaperPicks

Martin Jaggi

EPFL, School of Computer and Communication Sciences, Lausanne, Switzerland

24 papers at tracked venues · 23 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
    ACL 2026 ·
    Alejandro Hernández-Cano
  2. Attention with Markov: A Curious Case of Single-layer Transformers
  3. CoTFormer: A Chain of Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
  4. Effective Interplay between Sparsity and Quantization: From Theory to Practice
  5. Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
  6. GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
  7. Improving Stochastic Cubic Newton with Momentum
  8. Intrinsic User-Centric Interpretability through Global Mixture of Experts
  9. On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
  10. Towards Fully FP8 GEMM LLM Training at Scale
  11. URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
  12. Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
  13. CoBo: Collaborative Learning via Bilevel Optimization
  14. DOGE: Domain Reweighting with Generalization Estimation
  15. DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
  16. Ghost Noise for Regularizing Deep Neural Networks
  17. LASER: Linear Compression in Wireless Distributed Optimization
  18. Layer-wise linear mode connectivity
  19. On Convergence of Incremental Gradient for Non-convex Smooth Functions
  20. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
  21. Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
  22. Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
  23. Spectral Preconditioning for Gradient Methods on Graded Non-convex Functions
  24. The Privacy Power of Correlated Noise in Decentralized Learning