PPaperPicks

Zhao Song

University of California Berkeley, CA, USA

52 papers at tracked venues · 33 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
  2. A Fast Optimization View: Reformulating Single Layer Attention in LLM Based on Tensor and SVM Trick, and Solving It in Matrix Multiplication Time
  3. An Iterative Algorithm for Rescaled Hyperbolic Functions Regression
  4. Attention Mechanism, Max-Affine Partition, and Universal Approximation
  5. Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
  6. Binary Hypothesis Testing for Softmax Models and Leverage Score Models
  7. Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent
  8. Circuit Complexity Bounds for RoPE-based Transformer Architecture
  9. Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
  10. Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
  11. Deterministic Sparse Fourier Transform for Continuous Signals with Frequency Gap
  12. Differential Privacy Mechanisms in Neural Tangent Kernel Regression
  13. Differential Privacy for Euclidean Jordan Algebra with Applications to Private Symmetric Cone Programming
    NeurIPS 2025 · Zhao Song
  14. Discrepancy Minimization in Input-Sparsity Time
  15. Dissecting Submission Limit in Desk-Rejections: A Mathematical Analysis of Fairness in AI Conference Policies
  16. Dynamic Maintenance of Kernel Density Estimation Data Structure: From Practice to Theory
  17. Efficient $k$-Sparse Band-Limited Interpolation with Improved Approximation Ratio
  18. Efficient Alternating Minimization with Applications to Weighted Low Rank Approximation
    ICLR 2025 · Zhao Song
  19. Fast Sampling for Privacy-Preserving Lazy Multiplicative Weight Update
  20. Faster Algorithms for Structured John Ellipsoid Computation
  21. Faster Algorithms for Structured Linear and Kernel Support Vector Machines
  22. Force Matching with Relativistic Constraints: A Physics-Inspired Approach to Stable and Efficient Generative Modeling
  23. Fourier Circuits in Neural Networks and Transformers: A Case Study of Modular Arithmetic with Multiple Inputs
  24. Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency
  25. Fundamental Limits of Visual Autoregressive Transformers: Universal Approximation Abilities
  26. High-Order Flow Matching: Unified Framework and Sharp Statistical Rates
  27. In-Context Deep Learning via Transformer Models
  28. LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
  29. Looped ReLU MLPs May Be All You Need as Practical Programmable Computers
  30. NRFlow: Towards Noise-Robust Generative Modeling via High-Order Mechanism
  31. Numerical Pruning for Efficient Autoregressive Models
  32. On Differential Privacy for Adaptively Solving Search Problems via Sketching
  33. The Expressibility of Polynomial based Attention Scheme
    SIGKDD 2025 · Zhao Song
  34. Towards Infinite-Long Prefix in Transformer
  35. Unraveling the Smoothness Properties of Diffusion Models: A Gaussian Mixture Perspective
  36. When Can We Solve the Weighted Low Rank Approximation Problem in Truly Subquadratic Time?
  37. A General Algorithm for Solving Rank-one Matrix Sensing
  38. A Sublinear Adversarial Training Algorithm
  39. Algorithm and Hardness for Dynamic Attention Maintenance in Large Language Models
  40. Fast Dynamic Sampling for Determinantal Point Processes
    AISTATS 2024 · Zhao Song
  41. How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation
  42. How to Protect Copyright Data in Optimization of Large Language Models?
  43. Log-concave Sampling from a Convex Body with a Barrier: a Robust and Unified Dikin Walk
  44. Low Rank Matrix Completion via Robust Alternating Minimization in Nearly Linear Time
  45. Metric Transforms and Low Rank Representations of Kernels for Fast Attention
  46. On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity Analysis
  47. On Convergence of Federated Averaging Langevin Dynamics
  48. On Socially Fair Low-Rank Approximation and Column Subset Selection
    NeurIPS 2024 · Zhao Song
  49. On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)
  50. Solving Attention Kernel Regression Problem via Pre-conditioner
    AISTATS 2024 · Zhao Song
  51. The Closeness of In-Context Learning and Weight Shifting for Softmax Regression
  52. The Fine-Grained Complexity of Gradient Computation for Training Large Language Models