PPaperPicks

Minghui Fang

Zhejiang University, China

19 papers at tracked venues · 14 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling
    ACL 2025 · Minghui Fang
  2. ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
  3. Enhancing Multimodal Unified Representations for Cross Modal Generalization
  4. GTA: Towards Generative Text-To-Audio Retrieval via Multi-Scale Tokenizer
    InterSpeech 2025 · Minghui Fang
  5. Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
  6. MelRe: Vision-Based Mel-Spectrogram Restoration
  7. Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
  8. Multimodal Conditional Retrieval with High Controllability
  9. OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
  10. Open-Set Cross Modal Generalization via Multimodal Unified Representation
  11. Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
  12. Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
  13. Speech Watermarking with Discrete Intermediate Representations
  14. Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
  15. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
  16. Zero-resource Hallucination Detection for Text Generation via Graph-based Contextual Knowledge Triples Modeling
  17. AudioVSR: Enhancing Video Speech Recognition with Audio Data
  18. MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
  19. SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning