PPaperPicks

Zhangyang Wang

University of Texas at Austin, Cockrell School of Engineering, TX, USA

84 papers at tracked venues · 74 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. FlowMorph: Revealing an Optimizable Flow Latent Space for Controlled Image Morphing
  2. Improving the Throughput of Diffusion-based Large Language Models via a Training-Free Confidence-Aware Calibration
  3. Oscillation Inversion: Training-Free Image and Video Enhancement Through Oscillated Latents in Large Flow Models
  4. 4K4DGen: Panoramic 4D Generation at 4K Resolution
  5. A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs
  6. APOLLO: SGD-like Memory, AdamW-level Performance
  7. Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention
  8. Copy or Not? Reference-Based Face Image Restoration with Fine Details
  9. Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
  10. Efficient Image Generation with Variadic Attention Heads
  11. Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields
  12. FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splatting
  13. From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
  14. Graph-KV: Breaking Sequence via Injecting Structural Biases into Large Language Models
  15. HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model Training
  16. HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion Models
  17. Know Where You're Uncertain When Planning with Multimodal Foundation Models: A Formal Framework
  18. LLaMaFlex: Many-in-one LLMs via Generalized Pruning and Weight Sharing
  19. Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
  20. MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
  21. On the Provable Separation of Scales in Maximal Update Parameterization
  22. On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention for Long-Context LLM Serving
  23. One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
  24. PIPA: Preference Alignment as Prior-Informed Statistical Estimation
  25. R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
  26. REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
  27. Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding
  28. SAS: Simulated Attention Score
  29. SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
  30. Scaling Up Parameter Generation: A Recurrent Diffusion Approach
  31. Sparse Transfer Learning Accelerates and Enhances Certified Robustness: A Comprehensive Study
  32. Steepest Descent Density Control for Compact 3D Gaussian Splatting
  33. SteinDreamer: Variance Reduction for Text-to-3D Score Distillation via Stein Identity
  34. StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text
  35. Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
  36. Transformers Provably Learn Two-Mixture of Linear Classification via Gradient Flow
  37. Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing
  38. AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
  39. Compressing LLMs: The Truth is Rarely Pure and Never Simple
  40. DGBD: Depth Guided Branched Diffusion for Comprehensive Controllability in Multi-View Generation
  41. DP-OPT: Make Large Language Model Your Privacy-Preserving Prompt Engineer
  42. Data Distillation Can Be Like Vodka: Distilling More Times For Better Quality
  43. Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
  44. Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion Models
  45. Doubly Robust Instance-Reweighted Adversarial Training
  46. DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
  47. Efficient-3Dim: Learning a Generalizable Single-image Novel-view Synthesizer in One Day
  48. Expressive Gaussian Human Avatars from Monocular RGB Video
  49. FSGS: Real-Time Few-Shot View Synthesis Using Gaussian Splatting
  50. Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
  51. Flextron: Many-in-One Flexible Large Language Model
  52. Forget-Me-Not: Learning to Forget in Text-to-Image Diffusion Models
  53. Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional Encoding
  54. GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
  55. Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
  56. Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
  57. LLM-PBE: Assessing Data Privacy in Large Language Models
  58. LLaGA: Large Language and Graph Assistant
  59. Large Spatial Model: End-to-end Unposed Images to Semantic 3D
  60. Latent 3D Graph Diffusion
  61. Lift3D: Zero-Shot Lifting of Any 2D Vision Model to 3D
  62. LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPS
  63. LoCoCo: Dropping In Convolutions for Long Context Compression
  64. MM3DGS SLAM: Multi-modal 3D Gaussian Splatting for SLAM Using Vision, Depth, and Inertial Measurements
  65. Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild
  66. NeRF as Pretraining at Scale: Generalizable 3D-Aware Semantic Representation Learning from View Prediction
  67. OpenBias: Open-Set Bias Detection in Text-to-Image Generative Models
  68. Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
  69. PAIR Diffusion: A Comprehensive Multimodal Object-Level Image Editor
  70. POPE: 6-DoF Promptable Pose Estimation of Any Object, in Any Scene, with One Reference
  71. Polynomial Width is Sufficient for Set Representation with High-dimensional Features
  72. Principled Architecture-aware Scaling of Hyperparameters
  73. Prompt-Free Diffusion: Taking "Text" Out of Text-to-Image Diffusion Models
  74. RankMean: Module-Level Importance Score for Merging Fine-tuned LLM Models
  75. Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
  76. Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A Benchmark
  77. Safe and Robust Watermark Injection with a Single OoD Image
  78. Social Reward: Evaluating and Enhancing Generative AI through Million-User Feedback from an Online Creative Community
  79. Sparse Cocktail: Every Sparse Pattern Every Sparse Ratio All At Once
  80. Taming Mode Collapse in Score Distillation for Text-to-3D Generation
  81. Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis
  82. Turning A Curse into A Blessing: Data-Aware Memory-Efficient Training of Graph Neural Networks by Dynamic Exiting
  83. VersatileGaussian: Real-Time Neural Rendering for Versatile Tasks Using Gaussian Splatting
  84. Zero-Painter: Training-Free Layout Control for Text-to-Image Synthesis