PPaperPicks

Xixin Wu

28 papers at tracked venues · 10 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
  2. Masked Text-to-Audio Flow-Matching and Reward Feedback Optimization
  3. UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
  4. ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
  5. Autoregressive Speech Synthesis without Vector Quantization
  6. Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains
  7. Defending Unauthorized Voice Cloning with Watermark-Aware Codecs
  8. DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
  9. EEG-based Speech Decoding Based on Multi-mode Joint Modeling
  10. Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
  11. Generate, Discriminate, Evolve: Enhancing Context Faithfulness via Fine-Grained Sentence-Level Self-Evolution
  12. Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
  13. On the Within-class Variation Issue in Alzheimer's Disease Detection
  14. RAG-Zeval: Enhancing RAG Responses Evaluator through End-to-End Reasoning and Ranking-Based Reinforcement Learning
  15. WAKE: Watermarking Audio with Key Enrichment
  16. Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers
  17. CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
  18. Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
  19. Large Language Model-based FMRI Encoding of Language Functions for Subjects with Neurocognitive Disorder
  20. Natural Language Embedded Programs for Hybrid Language Symbolic Reasoning
  21. Prompting Large Language Models with Mispronunciation Detection and Diagnosis Abilities
  22. Rethinking Machine Ethics - Can LLMs Perform Moral Reasoning through the Lens of Moral Theories?
  23. Seamless Language Expansion: Enhancing Multilingual Mastery in Self-Supervised Models
  24. SimCalib: Graph Neural Network Calibration Based on Similarity between Nodes
  25. SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
  26. Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
  27. UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task Learner
  28. UniAudio: Towards Universal Audio Generation with Large Language Models