PPaperPicks

Ziyue Jiang

Zhejiang University, Hangzhou, China

21 papers at tracked venues · 15 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
  2. DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
  3. BrainLoc: Brain Signal-Based Object Detection with Multi-modal Alignment
  4. Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
  5. Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
  6. Speech Watermarking with Discrete Intermediate Representations
  7. TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
  8. Versatile Framework for Song Generation with Prompt-based Control
  9. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
  10. AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
  11. FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency
  12. GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
  13. InstructSpeech: Following Speech Editing Instructions via Large Language Models
  14. MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
  15. Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners
  16. Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
    ICLR 2024 · Ziyue Jiang
  17. MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
  18. MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
  19. Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis
  20. TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
  21. VoiceTuner: Self-Supervised Pre-training and Efficient Fine-tuning For Voice Generation