PPaperPicks

Jiatong Shi

35 papers at tracked venues · 11 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
  2. Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
  3. Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
  4. ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
    NeurIPS 2025 · Jiatong Shi
  5. Bridging Speech and Singing: Multi-stage Speech-Prompted Singing Voice Conversion with Speaker Embedding Adaptation
  6. Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
  7. Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
  8. ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
  9. ESPnet-SpeechLM: An Open Speech Language Model Toolkit
  10. FoodPuzzle: Toward Developing Large Language Model Agents as Autonomous Flavor Scientists
  11. MIKU-PAL: An Automated and Standardized Multimodal Method for Speech Paralinguistic and Affect Labeling
  12. OpusLM: A Family of Open Unified Speech Language Models
  13. Scalable Spontaneous Speech Dataset (SSSD): Crowdsourcing Data Collection to Promote Dialogue Research
  14. The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
  15. Uni-VERSA: Versatile Speech Assessment with a Unified Network
    InterSpeech 2025 · Jiatong Shi
  16. VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
    NAACL 2025 · Jiatong Shi
  17. AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
  18. CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
  19. EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
  20. ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
  21. ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
    InterSpeech 2024 · Jiatong Shi
  22. MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
    InterSpeech 2024 · Jiatong Shi
  23. Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners
  24. Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
    ICLR 2024 · Jiatong Shi
  25. Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
  26. OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
  27. PL-TTS: A Generalizable Prompt-based Diffusion TTS Augmented by Large Language Model
  28. Self-supervised Speech Representations Still Struggle with African American Vernacular English
  29. SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models
  30. Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing
    InterSpeech 2024 · Jiatong Shi
  31. The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
  32. TokSing: Singing Voice Synthesis based on Discrete Tokens
  33. Towards Robust Speech Representation Learning for Thousands of Languages
  34. UniAudio: Towards Universal Audio Generation with Large Language Models
  35. Wav2Gloss: Generating Interlinear Glossed Text from Speech