PPaperPicks

Xize Cheng

32 papers at tracked venues · 29 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Masked Text-to-Audio Flow-Matching and Reward Feedback Optimization
  2. SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness
  3. A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
  4. AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language Models
    NeurIPS 2025 · Xize Cheng
  5. CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling
  6. ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
  7. GTA: Towards Generative Text-To-Audio Retrieval via Multi-Scale Tokenizer
  8. Multimodal Conditional Retrieval with High Controllability
  9. OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
  10. OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
    ICLR 2025 · Xize Cheng
  11. PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue
  12. Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
  13. SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language
  14. T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
  15. VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?
    ICLR 2025 · Xize Cheng
  16. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
  17. AudioLCM: Efficient and High-Quality Text-to-Audio Generation with Minimal Inference Steps
  18. AudioVSR: Enhancing Video Speech Recognition with Audio Data
  19. Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented Prompts
  20. Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
  21. Extending Multi-modal Contrastive Representations
  22. FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
  23. InstructSpeech: Following Speech Editing Instructions via Large Language Models
  24. MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
  25. Rethinking the Multimodal Correlation of Multimodal Sequential Learning via Generalizable Attentional Results Alignment
  26. SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
  27. SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning
  28. Text-to-Song: Towards Controllable Music Generation Incorporating Vocal and Accompaniment
  29. TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation
    ACL 2024 · Xize Cheng
  30. Uni-Dubbing: Zero-Shot Speech Synthesis from Visual Articulation
  31. VoiceTuner: Self-Supervised Pre-training and Efficient Fine-tuning For Voice Generation
  32. Wav2SQL: Direct Generalizable Speech-To-SQL Parsing