PPaperPicks

Lei Xie

Northwestern Polytechnical University, School of Computer Science, Xi'an, China

41 papers at tracked venues · 17 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Hearing More with Less: Multi-Modal Retrieval-and-Selection Augmented Conversational LLM-Based ASR
  2. KALL-E: Autoregressive Speech Synthesis with Next-Distribution Prediction
  3. LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
  4. WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
  5. WenetSpeech-Yue: A Large-Scale Cantonese Speech Corpus with Multi-dimensional Annotation
  6. CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-car Speech Separation with Distributed Heterogeneous Arrays
  7. Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
  8. Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
  9. Drop the Beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
  10. DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
  11. EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
  12. Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning
  13. FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
  14. Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
  15. GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
  16. LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
  17. Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
  18. MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
  19. Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty
  20. StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
  21. Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
  22. Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
  23. Tractography-Guided Dual-Label Collaborative Learning for Multi-Modal Cranial Nerves Parcellation
    ACM MM 2025 · Lei Xie
  24. U-SAM: An Audio Language Model for Unified Speech, Audio, and Music Understanding
  25. Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR
  26. A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
  27. AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
  28. BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
  29. DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion
  30. Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
  31. RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
  32. SCDNet: Self-supervised Learning Feature based Speaker Change Detection
  33. SEQ-former: A context-enhanced and efficient automatic speech recognition framework
  34. Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
  35. StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
  36. Text-aware and Context-aware Expressive Audiobook Speech Synthesis
  37. Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
  38. Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
  39. UniStyle: Unified Style Modeling for Speaking Style Captioning and Stylistic Speech Synthesis
  40. Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
  41. WenetSpeech4TTS: A 12, 800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark