PPaperPicks

Dongchao Yang

20 papers at tracked venues · 15 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
  2. Masked Text-to-Audio Flow-Matching and Reward Feedback Optimization
  3. UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
  4. ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
    ICML 2025 · Dongchao Yang
  5. ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
  6. DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
  7. InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
  8. MoonCast: High-Quality Zero-Shot Podcast Generation
  9. Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
  10. SpeechSEC: A Unified Multi-Task Framework for Speech Synthesis, Editing, and Continuation
  11. AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
  12. CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
  13. InstructSpeech: Following Speech Editing Instructions via Large Language Models
  14. Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners
  15. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
  16. PromptTTS 2: Describing and Generating Voices with Text Prompt
  17. SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
    InterSpeech 2024 · Dongchao Yang
  18. UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task Learner
    NeurIPS 2024 · Dongchao Yang
  19. UniAudio: Towards Universal Audio Generation with Large Language Models
    ICML 2024 · Dongchao Yang
  20. VoiceTuner: Self-Supervised Pre-training and Efficient Fine-tuning For Voice Generation