PPaperPicks

Zhou Zhao

Zhejiang University, College of Computer Science, Hangzhou, China

111 papers at tracked venues · 100 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
  2. Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
  3. DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
  4. Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
  5. Iterative Self-Correction for Text-Driven Person Re-Identification with Large Vision-Language Models
  6. Masked Text-to-Audio Flow-Matching and Reward Feedback Optimization
  7. Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-Speech
  8. SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness
  9. Unified Thinker: A General Reasoning Core for Image Generation
  10. View-R1: Asymmetric Policy Optimization for Difficulty-Aware Multimodal Reinforcement Learning
  11. VoxMind: An End-to-End Agentic Spoken Dialogue System
  12. WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
  13. A Multimodal Evaluation Framework for Spatial Audio Playback Systems: From Localization to Listener Preference
  14. AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language Models
  15. AnomalyCoT: A Multi-Scenario Chain-of-Thought Dataset for Multimodal Large Language Models
  16. Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
  17. CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling
  18. CodeSync: Synchronizing Large Language Models with Dynamic Code Evolution at Scale
  19. ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
  20. Data-Efficiently Learn Large Language Model for Universal 3D Scene Perception
  21. Dataflow-Guided Neuro-Symbolic Language Models for Type Inference
  22. EAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic Integration
  23. EcoFace: Audio-Visual Emotional Co-Disentanglement Speech-Driven 3D Talking Face Generation
  24. Enhancing Multimodal Unified Representations for Cross Modal Generalization
  25. ExpTalk: Diverse Emotional Expression via Adaptive Disentanglement and Refined Alignment for Speech-Driven 3D Facial Animation
  26. FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
  27. FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio Generation
  28. GTA: Towards Generative Text-To-Audio Retrieval via Multi-Scale Tokenizer
  29. GenSpace: Benchmarking Spatially-Aware Image Generation
  30. ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
  31. IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
  32. ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
  33. InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model
  34. Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
  35. MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
  36. MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations
  37. MelodyEdit: Zero-shot Music Editing with Disentangled Inversion Control
  38. MergeNet: Knowledge Migration Across Heterogeneous Models, Tasks, and Modalities
  39. Multimodal Conditional Retrieval with High Controllability
  40. Non-Natural Image Understanding with Advancing Frequency-based Vision Encoders
  41. OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use
  42. OmniAudio: Generating Spatial Audio from 360-Degree Video
  43. OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
  44. OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
  45. Open-Set Cross Modal Generalization via Multimodal Unified Representation
  46. Orient Anything V2: Unifying Orientation and Rotation Understanding
  47. Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
  48. RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
  49. Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
  50. RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
  51. SPMDM: Enhancing Masked Diffusion Models through Simplifying Sampling Path
  52. STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation
  53. Seeking and Updating with Live Visual Knowledge
  54. Sign2Vis: Automated Data Visualization from Sign Language
  55. SkinGEN: an Explainable Dermatology Diagnosis-to-Generation Framework with Interactive Vision-Language Models
  56. SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language
  57. Speech Watermarking with Discrete Intermediate Representations
  58. T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
  59. TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
  60. TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow Matching
  61. ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and Editing
  62. Towards Transformer-Based Aligned Generation with Self-Coherence Guidance
  63. Versatile Framework for Song Generation with Prompt-based Control
  64. Vinci: Deep Thinking in Text-to-Image Generation using Unified Model with Reinforcement Learning
  65. VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?
  66. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
  67. AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
  68. Action Imitation in Common Action Space for Customized Action Image Synthesis
  69. AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
  70. AudioLCM: Efficient and High-Quality Text-to-Audio Generation with Minimal Inference Steps
  71. Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented Prompts
  72. Calibrating Prompt from History for Continual Vision-Language Retrieval and Grounding
  73. Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
  74. Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
  75. Cross-modal Observation Hypothesis Inference
  76. E3: Exploring Embodied Emotion Through A Large-Scale Egocentric Video Dataset
  77. EAGER: Two-Stream Generative Recommender with Behavior-Semantic Collaboration
  78. Extending Multi-modal Contrastive Representations
  79. FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
  80. Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
  81. GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
  82. InstructSpeech: Following Speech Editing Instructions via Large Language Models
  83. Low-rank Prompt Interaction for Continual Vision-Language Retrieval
  84. MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer
  85. MPOD123: One Image to 3D Content Generation Using Mask-Enhanced Progressive Outline-to-Detail Optimization
  86. MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
  87. Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners
  88. Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
  89. MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
  90. MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
  91. MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
  92. Multimodal Pretraining, Adaptation, and Generation for Recommendation: A Survey
  93. Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
  94. Non-confusing Generation of Customized Concepts in Diffusion Models
  95. Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
  96. Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis
  97. Rethinking the Multimodal Correlation of Multimodal Sequential Learning via Generalizable Attentional Results Alignment
  98. Robust Singing Voice Transcription Serves Synthesis
  99. Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion
  100. Semantic Alignment for Multimodal Large Language Models
  101. Semantic Codebook Learning for Dynamic Recommendation Models
  102. Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
  103. StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
  104. SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning
  105. TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
  106. Text-to-Song: Towards Controllable Music Generation Incorporating Vocal and Accompaniment
  107. TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation
  108. Uni-Dubbing: Zero-Shot Speech Synthesis from Visual Articulation
  109. UniAudio: Towards Universal Audio Generation with Large Language Models
  110. VoiceTuner: Self-Supervised Pre-training and Efficient Fine-tuning For Voice Generation
  111. Wav2SQL: Direct Generalizable Speech-To-SQL Parsing