PPaperPicks

Yifan Peng

NVIDIA Corporation, Santa Clara, CA, USA

19 papers at tracked venues · 4 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. Context-aware Dynamic Pruning for Speech Foundation Models
  2. DYNAC: Dynamic Vocabulary-based Non-Autoregressive Contextualization for Speech Recognition
  3. ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
  4. ESPnet-SpeechLM: An Open Speech Language Model Toolkit
  5. Enhancing Audiovisual Speech Recognition Through Bifocal Preference Optimization
  6. Exploring Linear Variant Transformers and k-NN Memory Inference for Long-Form ASR
  7. Granary: Speech Recognition and Translation Dataset in 25 European Languages
  8. Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
  9. OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
  10. OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
    InterSpeech 2025 · Yifan Peng
  11. OpusLM: A Family of Open Unified Speech Language Models
  12. VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
    NAACL 2025 · Yifan Peng
  13. Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
  14. MULTI-CONVFORMER: Extending Conformer with Multiple Convolution Kernels
  15. OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
    InterSpeech 2024 · Yifan Peng
  16. OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
    ACL 2024 · Yifan Peng
  17. On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
  18. Towards Robust Speech Representation Learning for Thousands of Languages
  19. UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions