PPaperPicks

Shinji Watanabe

Carnegie Mellon University, Pittsburgh, PA, USA

76 papers at tracked venues · 18 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
  2. CSPB: Conversational Speech Processing Benchmark for Self-supervised Speech Models
  3. Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
  4. Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
  5. POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
  6. PRiSM: Benchmarking Phone Realization in Speech Models
  7. PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding
  8. Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
  9. AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language Models
  10. ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
  11. CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
  12. Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
  13. Context-Driven Dynamic Pruning for Large Speech Foundation Models
  14. Context-aware Dynamic Pruning for Speech Foundation Models
  15. DYNAC: Dynamic Vocabulary-based Non-Autoregressive Contextualization for Speech Recognition
  16. DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
  17. Differentiable K-means for Fully-optimized Discrete Token-based ASR
  18. ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
  19. ESPnet-SpeechLM: An Open Speech Language Model Toolkit
  20. Enhancing Audiovisual Speech Recognition Through Bifocal Preference Optimization
  21. Explainable Depression Detection using Masked Hard Instance Mining
  22. Exploring Linear Variant Transformers and k-NN Memory Inference for Long-Form ASR
  23. GALAXY: A Large-Scale Open-Domain Dataset for Multimodal Learning
  24. Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
  25. Interspeech 2025 URGENT Speech Enhancement Challenge
  26. Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
  27. Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
  28. OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
  29. OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
  30. On-device Streaming Discrete Speech Units
  31. OpusLM: A Family of Open Unified Speech Language Models
  32. Pick and Summarize: Integrating Extractive and Abstractive Speech Summarization
  33. Scalable Spontaneous Speech Dataset (SSSD): Crowdsourcing Data Collection to Promote Dialogue Research
  34. Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
  35. SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
  36. Summarizing Speech: A Comprehensive Survey
  37. Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
  38. The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
  39. The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
  40. The Text-to-speech in the Wild (TITW) Database
  41. Uni-VERSA: Versatile Speech Assessment with a Unified Network
  42. VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
  43. VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
  44. AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
  45. Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
  46. Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
  47. Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
  48. Convolution-Augmented Parameter-Efficient Fine-Tuning for Speech Recognition
  49. Cross-Talk Reduction
  50. Decoder-only Architecture for Streaming End-to-end Speech Recognition
  51. DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
  52. EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
  53. EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
  54. ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
  55. FastAdaSP: Multitask-Adapted Efficient Inference for Large Speech Language Model
  56. Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
  57. ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
  58. MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
  59. MULTI-CONVFORMER: Extending Conformer with Multiple Convolution Kernels
  60. Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
  61. Neural Blind Source Separation and Diarization for Distant Speech Recognition
  62. OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
  63. OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
  64. On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
  65. On the Evaluation of Speech Foundation Models for Spoken Language Understanding
  66. Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
  67. Self-Supervised Speech Representations are More Phonetic than Semantic
  68. Self-training ASR Guided by Unsupervised ASR Teacher
  69. Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing
  70. SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
  71. The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
  72. To what extent can ASV systems naturally defend against spoofing attacks?
  73. Towards Robust Speech Representation Learning for Thousands of Languages
  74. URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
  75. UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
  76. Wav2Gloss: Generating Interlinear Glossed Text from Speech