P
PaperPicks
Conferences
Xie Chen
Shanghai Jiao Tong University, China
50 papers at tracked venues · 30 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0000-0001-7423-617X ↗
Google Scholar ↗
Homepage ↗
Venues
InterSpeech
×16
ACL
×14
AAAI
×7
ACM MM
×5
EMNLP
×3
NeurIPS
×2
ICCV
×1
ICML
×1
IJCAI
×1
Frequent coauthors
Ziyang Ma
DBLP profile ↗
ORCID search ↗
×6
Wenxi Chen
DBLP profile ↗
ORCID search ↗
×3
Yifan Yang
DBLP profile ↗
ORCID search ↗
×3
Yiwei Guo
DBLP profile ↗
ORCID search ↗
×2
Tianrui Wang
DBLP profile ↗
ORCID search ↗
×2
Xiquan Li
DBLP profile ↗
ORCID search ↗
×2
Yakun Song
DBLP profile ↗
ORCID search ↗
×2
Guanrou Yang
DBLP profile ↗
ORCID search ↗
×2
Zheng Lian
DBLP profile ↗
ORCID search ↗
×2
Tao Liu
DBLP profile ↗
ORCID search ↗
×2
Chenyuan Zhang
DBLP profile ↗
ORCID search ↗
×1
Haitao Li
DBLP profile ↗
ORCID search ↗
×1
Papers
AHAMask: Reliable Task Specification for Large Audio Language Models Without Instructions
AAAI 2026
·
Yiwei Guo
DBLP profile ↗
ORCID search ↗
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
ACL 2026
·
Tianrui Wang
DBLP profile ↗
ORCID search ↗
FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining
ACL 2026
·
Xiquan Li
DBLP profile ↗
ORCID search ↗
Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
ACL 2026
·
Chenyuan Zhang
DBLP profile ↗
ORCID search ↗
MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
ACL 2026
·
Xiquan Li
DBLP profile ↗
ORCID search ↗
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
ACL 2026
·
Haitao Li
DBLP profile ↗
ORCID search ↗
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
ACL 2026
·
Wenxi Chen
DBLP profile ↗
ORCID search ↗
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
ACL 2026
·
Yifan Yang
DBLP profile ↗
ORCID search ↗
WaveEx: Accelerating Flow Matching-based Speech Generation via Wavelet-guided Extrapolation
AAAI 2026
·
Xiaoqian Liu
DBLP profile ↗
ORCID search ↗
Accelerating Diffusion-based Text-to-Speech Model Trainingwith Dual Modality Alignment
InterSpeech 2025
·
Jeongsoo Choi
DBLP profile ↗
ORCID search ↗
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
InterSpeech 2025
·
Qixi Zheng
DBLP profile ↗
ORCID search ↗
Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
ICCV 2025
·
Xiao Li
DBLP profile ↗
ORCID search ↗
ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering
AAAI 2025
·
Yakun Song
DBLP profile ↗
ORCID search ↗
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
ACM MM 2025
·
Guanrou Yang
DBLP profile ↗
ORCID search ↗
Empowering Large Language Models for End-to-End Speech Translation Leveraging Synthetic Data
InterSpeech 2025
·
Yu Pu
DBLP profile ↗
ORCID search ↗
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
EMNLP 2025
·
Pengchao Feng
DBLP profile ↗
ORCID search ↗
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
InterSpeech 2025
·
Mingyu Cui
DBLP profile ↗
ORCID search ↗
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
ACL 2025
·
Yushen Chen
DBLP profile ↗
ORCID search ↗
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
ACL 2025
·
Yifan Yang
DBLP profile ↗
ORCID search ↗
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
InterSpeech 2025
·
Yiwei Guo
DBLP profile ↗
ORCID search ↗
Language Model Can Listen While Speaking
AAAI 2025
·
Ziyang Ma
DBLP profile ↗
ORCID search ↗
MER 2025: When Affective Computing Meets Large Language Models
ACM MM 2025
·
Zheng Lian
DBLP profile ↗
ORCID search ↗
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
NeurIPS 2025
·
Ziyang Ma
DBLP profile ↗
ORCID search ↗
MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language Models
EMNLP 2025
·
Yuezhang Peng
DBLP profile ↗
ORCID search ↗
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
ACL 2025
·
Yexing Du
DBLP profile ↗
ORCID search ↗
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
ACM MM 2025
·
Yifan Yang
DBLP profile ↗
ORCID search ↗
SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
ACL 2025
·
Wenxi Chen
DBLP profile ↗
ORCID search ↗
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
ACL 2025
·
Keqi Deng
DBLP profile ↗
ORCID search ↗
Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration
AAAI 2025
·
Ziyang Ma
DBLP profile ↗
ORCID search ↗
Towards Reliable Large Audio Language Model
ACL 2025
·
Ziyang Ma
DBLP profile ↗
ORCID search ↗
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
EMNLP 2025
·
Ruiqi Yan
DBLP profile ↗
ORCID search ↗
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
InterSpeech 2025
·
Hanglei Zhang
DBLP profile ↗
ORCID search ↗
VQTalker: Towards Multilingual Talking Avatars Through Facial Motion Tokenization
AAAI 2025
·
Tao Liu
DBLP profile ↗
ORCID search ↗
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
InterSpeech 2025
·
Jianheng Zhuo
DBLP profile ↗
ORCID search ↗
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
NeurIPS 2025
·
Tianrui Wang
DBLP profile ↗
ORCID search ↗
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
ACM MM 2024
·
Tao Liu
DBLP profile ↗
ORCID search ↗
AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
InterSpeech 2024
·
Anbai Jiang
DBLP profile ↗
ORCID search ↗
BAT: Learning to Reason about Spatial Sounds with Large Language Models
ICML 2024
·
Zhisheng Zheng
DBLP profile ↗
ORCID search ↗
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
IJCAI 2024
·
Wenxi Chen
DBLP profile ↗
ORCID search ↗
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
InterSpeech 2024
·
Ziyang Ma
DBLP profile ↗
ORCID search ↗
Improved Factorized Neural Transducer Model For Text-only Domain Adaptation
InterSpeech 2024
·
Junzhe Liu
DBLP profile ↗
ORCID search ↗
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
InterSpeech 2024
·
Peng Wang
DBLP profile ↗
ORCID search ↗
LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
InterSpeech 2024
·
Zheshu Song
DBLP profile ↗
ORCID search ↗
MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
ACM MM 2024
·
Zheng Lian
DBLP profile ↗
ORCID search ↗
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
InterSpeech 2024
·
Guanrou Yang
DBLP profile ↗
ORCID search ↗
On the Effectiveness of Acoustic BPE in Decoder-Only TTS
InterSpeech 2024
·
Bohan Li
DBLP profile ↗
ORCID search ↗
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
InterSpeech 2024
·
Yakun Song
DBLP profile ↗
ORCID search ↗
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
InterSpeech 2024
·
Xuankai Chang
DBLP profile ↗
ORCID search ↗
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
AAAI 2024
·
Chenpeng Du
DBLP profile ↗
ORCID search ↗
emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
ACL 2024
·
Ziyang Ma
DBLP profile ↗
ORCID search ↗