PPaperPicks

Di Wang

12 papers at tracked venues · 10 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Reinforcement Learning on Pre-Training Data
  2. TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model
  3. Union-of-Experts: Neurons in Mixture-of-Experts are Secretly Routers
  4. Autonomy-of-Experts Models
  5. DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
  6. Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs
  7. HMoE: Heterogeneous Mixture of Experts for Language Modeling
  8. Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment
  9. Scaling Laws for Floating-Point Quantization Training
  10. The Security Threat of Compressed Projectors in Large Vision-Language Models
  11. Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling
  12. Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning