PPaperPicks

Shengpeng Ji

29 papers at tracked venues · 26 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
  2. Dual-Reasoner: Bridging Interleaved Atomicity and Streaming Latency via Thinking-while-Talking
  3. VoxMind: An End-to-End Agentic Spoken Dialogue System
  4. WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
  5. AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language Models
  6. CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling
  7. ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
    ACL 2025 · Shengpeng Ji
  8. Enhancing Multimodal Unified Representations for Cross Modal Generalization
  9. GTA: Towards Generative Text-To-Audio Retrieval via Multi-Scale Tokenizer
  10. IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
  11. InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
  12. InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model
  13. Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
    ACL 2025 · Shengpeng Ji
  14. OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
  15. OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
  16. Open-Set Cross Modal Generalization via Multimodal Unified Representation
  17. Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
  18. SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language
  19. Speech Watermarking with Discrete Intermediate Representations
    AAAI 2025 · Shengpeng Ji
  20. T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
  21. TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
  22. UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
  23. VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?
  24. WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
  25. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
    ICLR 2025 · Shengpeng Ji
  26. AudioVSR: Enhancing Video Speech Recognition with Audio Data
  27. Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
  28. MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
    ACL 2024 · Shengpeng Ji
  29. SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning