P
PaperPicks
Conferences
Zihan Qiu
13 papers at tracked venues · 9 at CORE A* · active 2024–2025
DBLP profile ↗
ORCID search ↗
Venues
ICLR
×3
NeurIPS
×3
ACL
×2
NAACL
×2
AAAI
×1
EMNLP
×1
WSDM
×1
Frequent coauthors
Wenyu Du
DBLP profile ↗
ORCID search ↗
×2
Ka Man Lo
DBLP profile ↗
ORCID search ↗
×1
Yijun Yang
DBLP profile ↗
ORCID search ↗
×1
Zeyu Huang
DBLP profile ↗
ORCID search ↗
×1
Yongliang Wu
DBLP profile ↗
ORCID search ↗
×1
Hao Zhao
DBLP profile ↗
ORCID search ↗
×1
Papers
A Closer Look into Mixture-of-Experts in Large Language Models
NAACL 2025
·
Ka Man Lo
DBLP profile ↗
ORCID search ↗
A Controllable Examination for Long-Context Language Models
NeurIPS 2025
·
Yijun Yang
DBLP profile ↗
ORCID search ↗
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
ACL 2025
·
Zihan Qiu
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
NeurIPS 2025
·
Zihan Qiu
Layerwise Recurrent Router for Mixture-of-Experts
ICLR 2025
·
Zihan Qiu
Neo-TKGC: Enhancing Temporal Knowledge Graph Completion with Integrated Node Weights and Future Information
WSDM 2025
·
Zihan Qiu
Post-hoc Reward Calibration: A Case Study on Length Bias
ICLR 2025
·
Zeyu Huang
DBLP profile ↗
ORCID search ↗
Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
AAAI 2025
·
Yongliang Wu
DBLP profile ↗
ORCID search ↗
Empirical Study on Updating Key-Value Memories in Transformer Feed-forward Layers
ICLR 2024
·
Zihan Qiu
HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts
ACL 2024
·
Hao Zhao
DBLP profile ↗
ORCID search ↗
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
NeurIPS 2024
·
Wenyu Du
DBLP profile ↗
ORCID search ↗
Unlocking Continual Learning Abilities in Language Models
EMNLP 2024
·
Wenyu Du
DBLP profile ↗
ORCID search ↗
Unlocking Emergent Modularity in Large Language Models
NAACL 2024
·
Zihan Qiu