PPaperPicks

Qian Chen

Alibaba Group, DAMO Academy, Speech Lab, China

25 papers at tracked venues · 22 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
  2. GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling
  3. Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
  4. SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
  5. UniVocal: Unified Speech-Singing Code-Switching Synthesis
  6. WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
  7. ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World
  8. CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
  9. ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
  10. EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
  11. Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party Conversation
  12. MelodyEdit: Zero-shot Music Editing with Disentangled Inversion Control
  13. Multimodal Fusion and Coherence Modeling for Video Topic Segmentation
  14. OmniAudio: Generating Spatial Audio from 360-Degree Video
  15. OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
  16. Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
  17. Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR Transcripts
  18. Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration
  19. Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
  20. ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and Editing
  21. UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
  22. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
  23. CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
  24. ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
  25. TruthReader: Towards Trustworthy Document Assistant Chatbot with Reliable Attribution