PPaperPicks

Zhanpeng Zhou

8 papers at tracked venues · 8 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
  2. On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
  3. On the Role of Label Noise in the Feature Learning Process
  4. Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late In Training
    ICLR 2025 · Zhanpeng Zhou
  5. The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
  6. Batch Normalization Is Blind to the First and Second Derivatives of the Loss
    AAAI 2024 · Zhanpeng Zhou
  7. Going Beyond Neural Network Feature Similarity: The Network Feature Complexity and Its Interpretation Using Category Theory
  8. On the Emergence of Cross-Task Linearity in Pretraining-Finetuning Paradigm
    ICML 2024 · Zhanpeng Zhou