PPaperPicks

Zhen Ye

11 papers at tracked venues · 7 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Inference-time Scaling for Diffusion-based Audio Super-resolution
  2. AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness
  3. Boosting Policy and Process Reward Models with Monte Carlo Tree Search in Open-Domain QA
  4. Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
  5. Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
    AAAI 2025 · Zhen Ye
  6. LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
  7. ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges
  8. UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
  9. FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
  10. FlashSpeech: Efficient Zero-Shot Speech Synthesis
    ACM MM 2024 · Zhen Ye
  11. PyramidCodec: Hierarchical Codec for Long-form Music Generation in Audio Domain