PPaperPicks

Xie Chen

Shanghai Jiao Tong University, China

50 papers at tracked venues · 30 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AHAMask: Reliable Task Specification for Large Audio Language Models Without Instructions
  2. Evaluating the Expressive Appropriateness of Speech in Rich Contexts
  3. FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining
  4. Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
  5. MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
  6. ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
  7. SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
  8. Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
  9. WaveEx: Accelerating Flow Matching-based Speech Generation via Wavelet-guided Extrapolation
  10. Accelerating Diffusion-based Text-to-Speech Model Trainingwith Dual Modality Alignment
  11. Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
  12. Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
  13. ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering
  14. EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
  15. Empowering Large Language Models for End-to-End Speech Translation Leveraging Synthetic Data
  16. Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
  17. Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
  18. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
  19. GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
  20. LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
  21. Language Model Can Listen While Speaking
  22. MER 2025: When Affective Computing Meets Large Language Models
  23. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
  24. MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language Models
  25. Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
  26. Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
  27. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
  28. SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
  29. Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration
  30. Towards Reliable Large Audio Language Model
  31. URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
  32. Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
  33. VQTalker: Towards Multilingual Talking Avatars Through Facial Motion Tokenization
  34. VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
  35. Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
  36. AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
  37. AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
  38. BAT: Learning to Reason about Spatial Sounds with Large Language Models
  39. EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
  40. EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
  41. Improved Factorized Neural Transducer Model For Text-only Domain Adaptation
  42. Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
  43. LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
  44. MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
  45. MaLa-ASR: Multimedia-Assisted LLM-Based ASR
  46. On the Effectiveness of Acoustic BPE in Decoder-Only TTS
  47. TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
  48. The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
  49. UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
  50. emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation