P
PaperPicks
Conferences
Joon Son Chung
28 papers at tracked venues · 17 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0000-0001-7741-7275 ↗
Homepage ↗
Venues
InterSpeech
×10
CVPR
×5
NeurIPS
×3
AAAI
×2
ACM MM
×2
ICLR
×2
ACL
×1
EMNLP
×1
ICCV
×1
ICML
×1
Frequent coauthors
Chaeyoung Jung
DBLP profile ↗
ORCID search ↗
×3
Jeongsoo Choi
DBLP profile ↗
ORCID search ↗
×3
Dawit Mureja Argaw
DBLP profile ↗
ORCID search ↗
×3
Sung-Bin Kim
DBLP profile ↗
ORCID search ↗
×2
Ji-Hoon Kim
DBLP profile ↗
ORCID search ↗
×2
Kihyun Nam
DBLP profile ↗
ORCID search ↗
×2
Jee-weon Jung
DBLP profile ↗
ORCID search ↗
×2
Sonal Kumar
DBLP profile ↗
ORCID search ↗
×1
Sungnyun Kim
DBLP profile ↗
ORCID search ↗
×1
Kang Zhang
DBLP profile ↗
ORCID search ↗
×1
Hyeonggon Ryu
DBLP profile ↗
ORCID search ↗
×1
Chenshuang Zhang
DBLP profile ↗
ORCID search ↗
×1
Papers
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
AAAI 2026
·
Sonal Kumar
DBLP profile ↗
ORCID search ↗
Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
ACL 2026
·
Sungnyun Kim
DBLP profile ↗
ORCID search ↗
AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
NeurIPS 2025
·
Chaeyoung Jung
DBLP profile ↗
ORCID search ↗
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
ICLR 2025
·
Sung-Bin Kim
DBLP profile ↗
ORCID search ↗
Accelerating Diffusion-based Text-to-Speech Model Trainingwith Dual Modality Alignment
InterSpeech 2025
·
Jeongsoo Choi
DBLP profile ↗
ORCID search ↗
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
ACM MM 2025
·
Jeongsoo Choi
DBLP profile ↗
ORCID search ↗
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
EMNLP 2025
·
Jeongsoo Choi
DBLP profile ↗
ORCID search ↗
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
CVPR 2025
·
Ji-Hoon Kim
DBLP profile ↗
ORCID search ↗
High-Quality Joint Image and Video Tokenization with Causal VAE
ICLR 2025
·
Dawit Mureja Argaw
DBLP profile ↗
ORCID search ↗
InfiniteAudio: Infinite-Length Audio Generation with Consistency
InterSpeech 2025
·
Chaeyoung Jung
DBLP profile ↗
ORCID search ↗
Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
NeurIPS 2025
·
Kang Zhang
DBLP profile ↗
ORCID search ↗
SEED: Speaker Embedding Enhancement Diffusion Model
InterSpeech 2025
·
Kihyun Nam
DBLP profile ↗
ORCID search ↗
Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes
CVPR 2025
·
Hyeonggon Ryu
DBLP profile ↗
ORCID search ↗
The Text-to-speech in the Wild (TITW) Database
InterSpeech 2025
·
Jee-weon Jung
DBLP profile ↗
ORCID search ↗
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
NeurIPS 2025
·
Chenshuang Zhang
DBLP profile ↗
ORCID search ↗
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
ICCV 2025
·
Sung-Bin Kim
DBLP profile ↗
ORCID search ↗
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
InterSpeech 2024
·
Kihyun Nam
DBLP profile ↗
ORCID search ↗
ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions
InterSpeech 2024
·
Jiu Feng
DBLP profile ↗
ORCID search ↗
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
ICML 2024
·
Jongsuk Kim
DBLP profile ↗
ORCID search ↗
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
CVPR 2024
·
Youngjoon Jang
DBLP profile ↗
ORCID search ↗
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
InterSpeech 2024
·
Chaeyoung Jung
DBLP profile ↗
ORCID search ↗
Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding
ACM MM 2024
·
Jongbhin Woo
DBLP profile ↗
ORCID search ↗
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
AAAI 2024
·
Ji-Hoon Kim
DBLP profile ↗
ORCID search ↗
Lightweight Audio Segmentation for Long-form Speech Translation
InterSpeech 2024
·
Jaesong Lee
DBLP profile ↗
ORCID search ↗
Scaling Up Video Summarization Pretraining with Large Language Models
CVPR 2024
·
Dawit Mureja Argaw
DBLP profile ↗
ORCID search ↗
To what extent can ASV systems naturally defend against spoofing attacks?
InterSpeech 2024
·
Jee-weon Jung
DBLP profile ↗
ORCID search ↗
Towards Automated Movie Trailer Generation
CVPR 2024
·
Dawit Mureja Argaw
DBLP profile ↗
ORCID search ↗
VoxSim: A perceptual voice similarity dataset
InterSpeech 2024
·
Junseok Ahn
DBLP profile ↗
ORCID search ↗