P
PaperPicks
Conferences
Yifan Yang
Xiaomi Corp., Beijing, China
16 papers at tracked venues · 11 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0009-0003-0588-1812 ↗
Venues
ACL
×6
InterSpeech
×5
ACM MM
×3
AAAI
×1
ICLR
×1
Frequent coauthors
Hui Wang
DBLP profile ↗
ORCID search ↗
×2
Tianrui Wang
DBLP profile ↗
ORCID search ↗
×1
Guanrou Yang
DBLP profile ↗
ORCID search ↗
×1
Mingyu Cui
DBLP profile ↗
ORCID search ↗
×1
Yexing Du
DBLP profile ↗
ORCID search ↗
×1
Wenxi Chen
DBLP profile ↗
ORCID search ↗
×1
Ziyang Ma
DBLP profile ↗
ORCID search ↗
×1
Jianheng Zhuo
DBLP profile ↗
ORCID search ↗
×1
Peng Wang
DBLP profile ↗
ORCID search ↗
×1
Zengrui Jin
DBLP profile ↗
ORCID search ↗
×1
Zheshu Song
DBLP profile ↗
ORCID search ↗
×1
Zengwei Yao
DBLP profile ↗
ORCID search ↗
×1
Papers
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
ACL 2026
·
Tianrui Wang
DBLP profile ↗
ORCID search ↗
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
ACL 2026
·
Hui Wang
DBLP profile ↗
ORCID search ↗
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
ACL 2026
·
Yifan Yang
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
ACM MM 2025
·
Guanrou Yang
DBLP profile ↗
ORCID search ↗
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
InterSpeech 2025
·
Mingyu Cui
DBLP profile ↗
ORCID search ↗
FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
ACM MM 2025
·
Hui Wang
DBLP profile ↗
ORCID search ↗
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
ACL 2025
·
Yifan Yang
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
ACL 2025
·
Yexing Du
DBLP profile ↗
ORCID search ↗
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
ACM MM 2025
·
Yifan Yang
SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
ACL 2025
·
Wenxi Chen
DBLP profile ↗
ORCID search ↗
Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration
AAAI 2025
·
Ziyang Ma
DBLP profile ↗
ORCID search ↗
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
InterSpeech 2025
·
Jianheng Zhuo
DBLP profile ↗
ORCID search ↗
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
InterSpeech 2024
·
Peng Wang
DBLP profile ↗
ORCID search ↗
LibriheavyMix: A 20, 000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
InterSpeech 2024
·
Zengrui Jin
DBLP profile ↗
ORCID search ↗
LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
InterSpeech 2024
·
Zheshu Song
DBLP profile ↗
ORCID search ↗
Zipformer: A faster and better encoder for automatic speech recognition
ICLR 2024
·
Zengwei Yao
DBLP profile ↗
ORCID search ↗