PPaperPicks

Joon Son Chung

28 papers at tracked venues · 17 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
  2. Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
  3. AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
  4. AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
  5. Accelerating Diffusion-based Text-to-Speech Model Trainingwith Dual Modality Alignment
  6. AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
  7. Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
  8. From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
  9. High-Quality Joint Image and Video Tokenization with Causal VAE
  10. InfiniteAudio: Infinite-Length Audio Generation with Consistency
  11. Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
  12. SEED: Speaker Embedding Enhancement Diffusion Model
  13. Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes
  14. The Text-to-speech in the Wild (TITW) Database
  15. Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
  16. VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
  17. Disentangled Representation Learning for Environment-agnostic Speaker Recognition
  18. ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions
  19. EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
  20. Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
  21. FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
  22. Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding
  23. Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
  24. Lightweight Audio Segmentation for Long-form Speech Translation
  25. Scaling Up Video Summarization Pretraining with Large Language Models
  26. To what extent can ASV systems naturally defend against spoofing attacks?
  27. Towards Automated Movie Trailer Generation
  28. VoxSim: A perceptual voice similarity dataset