PPaperPicks

Ziyang Ma

Shanghai Jiao Tong University, Department of Computer Science and Engineering, AI Institute, MoE Key Lab of Artificial Intelligence, Shanghai, China

30 papers at tracked venues · 22 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Evaluating the Expressive Appropriateness of Speech in Rich Contexts
  2. FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining
  3. SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
  4. Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
  5. Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
  6. ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering
  7. EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
  8. Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
  9. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
  10. GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
  11. Language Model Can Listen While Speaking
    AAAI 2025 · Ziyang Ma
  12. MER 2025: When Affective Computing Meets Large Language Models
  13. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
    NeurIPS 2025 · Ziyang Ma
  14. Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
  15. MuPT: A Generative Symbolic Music Pretrained Transformer
  16. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
  17. Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration
    AAAI 2025 · Ziyang Ma
  18. Towards Reliable Large Audio Language Model
    ACL 2025 · Ziyang Ma
  19. URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
  20. VQTalker: Towards Multilingual Talking Avatars Through Facial Motion Tokenization
  21. Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
  22. BAT: Learning to Reason about Spatial Sounds with Large Language Models
  23. ChatMusician: Understanding and Generating Music Intrinsically with LLM
  24. EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
  25. EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
    InterSpeech 2024 · Ziyang Ma
  26. LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
  27. MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
  28. MaLa-ASR: Multimedia-Assisted LLM-Based ASR
  29. TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
  30. emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
    ACL 2024 · Ziyang Ma