PPaperPicks

Zhiyong Wu

Tsinghua University, Joint Research Center for Media Sciences, Beijing, China

30 papers at tracked venues · 19 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
  2. Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
  3. Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies
  4. UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
  5. A Dual-Branch Ensemble Framework for Personality Recognition Based on Multimodal Emotion Features
  6. DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
  7. DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
  8. Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
  9. HarmoniVox: Painting Voices to Match the Avatar's Soul
  10. LeVo: High-Quality Song Generation with Multi-Preference Alignment
  11. MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative Refinement
  12. MuCodec: Ultra Low-Bitrate Music Codec for Music Generation
  13. RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction
  14. StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
  15. VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweening
  16. WAKE: Watermarking Audio with Key Enrichment
  17. Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
  18. CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
  19. Comparing Discrete and Continuous Space LLMs for Speech Recognition
  20. Explore 3D Dance Generation via Reward Model from Automatically-Ranked Demonstrations
  21. LoRA-MER: Low-Rank Adaptation of Pre-Trained Speech Models for Multimodal Emotion Recognition Using Mutual Information
  22. Multimodal Emotion Captioning Using Large Language Model with Prompt Engineering
  23. Robust Representation Learning for Multimodal Emotion Recognition with Contrastive Learning and Mixup
  24. SECap: Speech Emotion Captioning with Large Language Model
  25. SimCalib: Graph Neural Network Calibration Based on Similarity between Nodes
  26. SongCreator: Lyrics-based Universal Song Generation
  27. Speaker Change Detection with Weighted-sum Knowledge Distillation based on Self-supervised Pre-trained Models
  28. SpeechCraft: A Fine-Grained Expressive Speech Dataset with Natural Language Description
  29. Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
  30. VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling