PPaperPicks

Shuai Wang

Chinese University of Hong Kong-Shenzhen (CUKH-SZ), Shenzhen Research Institute of Big Data, Shenzhen, China

16 papers at tracked venues · 7 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AHAMask: Reliable Task Specification for Large Audio Language Models Without Instructions
  2. USE: A Unified Model for Universal Sound Separation and Extraction
  3. Drop the Beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
  4. Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
  5. LeVo: High-Quality Song Generation with Multi-Preference Alignment
  6. PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
  7. REAL-T: Real Conversational Mixtures for Target Speaker Extraction
  8. SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
  9. SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
  10. TVC-MusicGen: Time-Varying Structure Control for Background Music Generation via Self-Supervised Training
  11. DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion
  12. On the Effectiveness of Acoustic BPE in Decoder-Only TTS
  13. UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
  14. WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
    InterSpeech 2024 · Shuai Wang
  15. WenetSpeech4TTS: A 12, 800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
  16. Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models