P
PaperPicks
Conferences
Jiatong Shi
35 papers at tracked venues · 11 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
InterSpeech
×19
ACL
×4
NAACL
×3
ICLR
×2
AAAI
×1
ACM MM
×1
EACL
×1
EMNLP
×1
ICML
×1
NeurIPS
×1
SIGKDD
×1
Frequent coauthors
Siddhant Arora
DBLP profile ↗
ORCID search ↗
×3
Jinchuan Tian
DBLP profile ↗
ORCID search ↗
×2
William Chen
DBLP profile ↗
ORCID search ↗
×2
Rongjie Huang
DBLP profile ↗
ORCID search ↗
×2
Yuning Wu
DBLP profile ↗
ORCID search ↗
×2
Haoran Wang
DBLP profile ↗
ORCID search ↗
×1
Guan-Ting Lin
DBLP profile ↗
ORCID search ↗
×1
Mingda Liu
DBLP profile ↗
ORCID search ↗
×1
Chien-yu Huang
DBLP profile ↗
ORCID search ↗
×1
Tenghao Huang
DBLP profile ↗
ORCID search ↗
×1
Yifan Cheng
DBLP profile ↗
ORCID search ↗
×1
Zaid Sheikh
DBLP profile ↗
ORCID search ↗
×1
Papers
BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
EACL 2026
·
Haoran Wang
DBLP profile ↗
ORCID search ↗
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
ACL 2026
·
Guan-Ting Lin
DBLP profile ↗
ORCID search ↗
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
ACL 2026
·
Siddhant Arora
DBLP profile ↗
ORCID search ↗
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
NeurIPS 2025
·
Jiatong Shi
Bridging Speech and Singing: Multi-stage Speech-Prompted Singing Voice Conversion with Speaker Embedding Adaptation
InterSpeech 2025
·
Mingda Liu
DBLP profile ↗
ORCID search ↗
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
InterSpeech 2025
·
Siddhant Arora
DBLP profile ↗
ORCID search ↗
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
ICLR 2025
·
Chien-yu Huang
DBLP profile ↗
ORCID search ↗
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
NAACL 2025
·
Siddhant Arora
DBLP profile ↗
ORCID search ↗
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
NAACL 2025
·
Jinchuan Tian
DBLP profile ↗
ORCID search ↗
FoodPuzzle: Toward Developing Large Language Model Agents as Autonomous Flavor Scientists
SIGKDD 2025
·
Tenghao Huang
DBLP profile ↗
ORCID search ↗
MIKU-PAL: An Automated and Standardized Multimodal Method for Speech Paralinguistic and Affect Labeling
InterSpeech 2025
·
Yifan Cheng
DBLP profile ↗
ORCID search ↗
OpusLM: A Family of Open Unified Speech Language Models
InterSpeech 2025
·
Jinchuan Tian
DBLP profile ↗
ORCID search ↗
Scalable Spontaneous Speech Dataset (SSSD): Crowdsourcing Data Collection to Promote Dialogue Research
InterSpeech 2025
·
Zaid Sheikh
DBLP profile ↗
ORCID search ↗
The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
InterSpeech 2025
·
William Chen
DBLP profile ↗
ORCID search ↗
Uni-VERSA: Versatile Speech Assessment with a Unified Network
InterSpeech 2025
·
Jiatong Shi
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
NAACL 2025
·
Jiatong Shi
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
AAAI 2024
·
Rongjie Huang
DBLP profile ↗
ORCID search ↗
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
InterSpeech 2024
·
Yongyi Zang
DBLP profile ↗
ORCID search ↗
EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
InterSpeech 2024
·
Tejes Srivastava
DBLP profile ↗
ORCID search ↗
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
InterSpeech 2024
·
Jee-weon Jung
DBLP profile ↗
ORCID search ↗
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
InterSpeech 2024
·
Jiatong Shi
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
InterSpeech 2024
·
Jiatong Shi
Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners
ACL 2024
·
Rongjie Huang
DBLP profile ↗
ORCID search ↗
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
ICLR 2024
·
Jiatong Shi
Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
ACM MM 2024
·
Yuning Wu
DBLP profile ↗
ORCID search ↗
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
InterSpeech 2024
·
Yifan Peng
DBLP profile ↗
ORCID search ↗
PL-TTS: A Generalizable Prompt-based Diffusion TTS Augmented by Large Language Model
InterSpeech 2024
·
Shuhua Li
DBLP profile ↗
ORCID search ↗
Self-supervised Speech Representations Still Struggle with African American Vernacular English
InterSpeech 2024
·
Kalvin Chang
DBLP profile ↗
ORCID search ↗
SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models
InterSpeech 2024
·
Yuxun Tang
DBLP profile ↗
ORCID search ↗
Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing
InterSpeech 2024
·
Jiatong Shi
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
InterSpeech 2024
·
Xuankai Chang
DBLP profile ↗
ORCID search ↗
TokSing: Singing Voice Synthesis based on Discrete Tokens
InterSpeech 2024
·
Yuning Wu
DBLP profile ↗
ORCID search ↗
Towards Robust Speech Representation Learning for Thousands of Languages
EMNLP 2024
·
William Chen
DBLP profile ↗
ORCID search ↗
UniAudio: Towards Universal Audio Generation with Large Language Models
ICML 2024
·
Dongchao Yang
DBLP profile ↗
ORCID search ↗
Wav2Gloss: Generating Interlinear Glossed Text from Speech
ACL 2024
·
Taiqi He
DBLP profile ↗
ORCID search ↗