PPaperPicks

Yuxuan Wang

12 papers at tracked venues · 10 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
  2. Language Model Can Listen While Speaking
  3. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
  4. QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
  5. SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation
  6. Sounding that Object: Interactive Object-Aware Image to Audio Generation
  7. Towards Reliable Large Audio Language Model
  8. Can Large Language Models Understand Spatial Audio?
  9. InstructME: An Instruction Guided Music Edit Framework with Latent Diffusion Models
  10. PolyVoice: Language Models for Speech to Speech Translation
  11. SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
  12. video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models