PPaperPicks

Yifan Yang

Xiaomi Corp., Beijing, China

16 papers at tracked venues · 11 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Evaluating the Expressive Appropriateness of Speech in Rich Contexts
  2. SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
  3. Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
    ACL 2026 · Yifan Yang
  4. EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
  5. Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
  6. FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
  7. GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
    ACL 2025 · Yifan Yang
  8. Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
  9. Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
    ACM MM 2025 · Yifan Yang
  10. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
  11. Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration
  12. VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
  13. Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
  14. LibriheavyMix: A 20, 000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
  15. LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
  16. Zipformer: A faster and better encoder for automatic speech recognition