P
PaperPicks
Conferences
Dongchao Yang
20 papers at tracked venues · 15 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
ACL
×5
ICML
×4
InterSpeech
×4
AAAI
×2
NeurIPS
×2
ACM MM
×1
EMNLP
×1
ICLR
×1
Frequent coauthors
Rongjie Huang
DBLP profile ↗
ORCID search ↗
×5
Yuanyuan Wang
DBLP profile ↗
ORCID search ↗
×2
Xueyuan Chen
DBLP profile ↗
ORCID search ↗
×2
Dingdong Wang
DBLP profile ↗
ORCID search ↗
×2
Zeqian Ju
DBLP profile ↗
ORCID search ↗
×2
Yuguo Yin
DBLP profile ↗
ORCID search ↗
×1
Liming Liang
DBLP profile ↗
ORCID search ↗
×1
Yichong Leng
DBLP profile ↗
ORCID search ↗
×1
Papers
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
AAAI 2026
·
Yuanyuan Wang
DBLP profile ↗
ORCID search ↗
Masked Text-to-Audio Flow-Matching and Reward Feedback Optimization
ACL 2026
·
Rongjie Huang
DBLP profile ↗
ORCID search ↗
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
ACL 2026
·
Yuanyuan Wang
DBLP profile ↗
ORCID search ↗
ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
ICML 2025
·
Dongchao Yang
ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
ACL 2025
·
Yuguo Yin
DBLP profile ↗
ORCID search ↗
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
InterSpeech 2025
·
Xueyuan Chen
DBLP profile ↗
ORCID search ↗
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
ACL 2025
·
Dingdong Wang
DBLP profile ↗
ORCID search ↗
MoonCast: High-Quality Zero-Shot Podcast Generation
NeurIPS 2025
·
Zeqian Ju
DBLP profile ↗
ORCID search ↗
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
EMNLP 2025
·
Dingdong Wang
DBLP profile ↗
ORCID search ↗
SpeechSEC: A Unified Multi-Task Framework for Speech Synthesis, Editing, and Continuation
InterSpeech 2025
·
Liming Liang
DBLP profile ↗
ORCID search ↗
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
AAAI 2024
·
Rongjie Huang
DBLP profile ↗
ORCID search ↗
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
InterSpeech 2024
·
Xueyuan Chen
DBLP profile ↗
ORCID search ↗
InstructSpeech: Following Speech Editing Instructions via Large Language Models
ICML 2024
·
Rongjie Huang
DBLP profile ↗
ORCID search ↗
Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners
ACL 2024
·
Rongjie Huang
DBLP profile ↗
ORCID search ↗
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
ICML 2024
·
Zeqian Ju
DBLP profile ↗
ORCID search ↗
PromptTTS 2: Describing and Generating Voices with Text Prompt
ICLR 2024
·
Yichong Leng
DBLP profile ↗
ORCID search ↗
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
InterSpeech 2024
·
Dongchao Yang
UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task Learner
NeurIPS 2024
·
Dongchao Yang
UniAudio: Towards Universal Audio Generation with Large Language Models
ICML 2024
·
Dongchao Yang
VoiceTuner: Self-Supervised Pre-training and Efficient Fine-tuning For Voice Generation
ACM MM 2024
·
Rongjie Huang
DBLP profile ↗
ORCID search ↗