PPaperPicks

Shiwei Liu

University of Oxford, Mathematical Institute, UK

28 papers at tracked venues · 23 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
  2. AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
  3. Composable Interventions for Language Models
  4. From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
  5. GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
  6. LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning
  7. Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
  8. Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
  9. Outlier-weighed Layerwise Sampling for LLM Fine-tuning
  10. SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
  11. The Curse of Depth in Large Language Models
  12. Visual Prompting Upgrades Neural Network Sparsification: A Data-Model Perspective
  13. AdaMerging: Adaptive Model Merging for Multi-Task Learning
  14. Advancing Dynamic Sparse Training by Exploring Optimization Opportunities
  15. AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
  16. CaM: Cache Merging for Memory-efficient LLMs Inference
  17. Dynamic Data Pruning for Automatic Speech Recognition
  18. Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs
  19. E2ENet: Dynamic Sparse Feature Fusion for Accurate and Efficient 3D Medical Image Segmentation
  20. FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
  21. Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional Encoding
  22. Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning
  23. Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
  24. MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
  25. NeurRev: Train Better Sparse Neural Network Practically via Neuron Revitalization
  26. Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
  27. Q-Hitter: A Better Token Oracle for Efficient LLM Inference via Sparse-Quantized KV Cache
  28. Sparse Cocktail: Every Sparse Pattern Every Sparse Ratio All At Once