P
PaperPicks
Conferences
Xiaoda Yang
22 papers at tracked venues · 15 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID 0009-0002-7297-4536 ↗
Homepage ↗
Venues
ACM MM
×5
AAAI
×3
EMNLP
×3
ICLR
×3
InterSpeech
×3
ACL
×2
SIGKDD
×1
WACV
×1
WWW
×1
Frequent coauthors
Minghui Fang
DBLP profile ↗
ORCID search ↗
×2
Dongjie Fu
DBLP profile ↗
ORCID search ↗
×2
Hao Li
DBLP profile ↗
ORCID search ↗
×1
Donglin Huang
DBLP profile ↗
ORCID search ↗
×1
Jiaqi Duan
DBLP profile ↗
ORCID search ↗
×1
Xueyi Zhang
DBLP profile ↗
ORCID search ↗
×1
Weicai Yan
DBLP profile ↗
ORCID search ↗
×1
Minjie Hong
DBLP profile ↗
ORCID search ↗
×1
Sijing Li
DBLP profile ↗
ORCID search ↗
×1
Kaixuan Luan
DBLP profile ↗
ORCID search ↗
×1
Jialong Zuo
DBLP profile ↗
ORCID search ↗
×1
Wenrui Liu
DBLP profile ↗
ORCID search ↗
×1
Papers
SpatialLogic-Bench: A Diagnostic Benchmark for Task-Oriented Spatiotemporal Reasoning
AAAI 2026
·
Xiaoda Yang
Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
AAAI 2026
·
Hao Li
DBLP profile ↗
ORCID search ↗
VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
WACV 2026
·
Donglin Huang
DBLP profile ↗
ORCID search ↗
BrainLoc: Brain Signal-Based Object Detection with Multi-modal Alignment
EMNLP 2025
·
Jiaqi Duan
DBLP profile ↗
ORCID search ↗
CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling
ACL 2025
·
Minghui Fang
DBLP profile ↗
ORCID search ↗
Choose Your Expert: Uncertainty-Guided Expert Selection for Continual Deepfake Detection
ACM MM 2025
·
Xueyi Zhang
DBLP profile ↗
ORCID search ↗
Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
ICLR 2025
·
Weicai Yan
DBLP profile ↗
ORCID search ↗
EAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic Integration
WWW 2025
·
Minjie Hong
DBLP profile ↗
ORCID search ↗
EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
ACM MM 2025
·
Sijing Li
DBLP profile ↗
ORCID search ↗
GTA: Towards Generative Text-To-Audio Retrieval via Multi-Scale Tokenizer
InterSpeech 2025
·
Minghui Fang
DBLP profile ↗
ORCID search ↗
MelRe: Vision-Based Mel-Spectrogram Restoration
InterSpeech 2025
·
Kaixuan Luan
DBLP profile ↗
ORCID search ↗
Multimodal Conditional Retrieval with High Controllability
SIGKDD 2025
·
Xiaoda Yang
PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue
EMNLP 2025
·
Dongjie Fu
DBLP profile ↗
ORCID search ↗
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
ACL 2025
·
Jialong Zuo
DBLP profile ↗
ORCID search ↗
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
ACM MM 2025
·
Wenrui Liu
DBLP profile ↗
ORCID search ↗
Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection
AAAI 2025
·
Yuhang Ma
DBLP profile ↗
ORCID search ↗
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
InterSpeech 2025
·
Ruofan Hu
DBLP profile ↗
ORCID search ↗
VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?
ICLR 2025
·
Xize Cheng
DBLP profile ↗
ORCID search ↗
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
ICLR 2025
·
Shengpeng Ji
DBLP profile ↗
ORCID search ↗
AudioVSR: Enhancing Video Speech Recognition with Audio Data
EMNLP 2024
·
Xiaoda Yang
Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented Prompts
ACM MM 2024
·
Dongjie Fu
DBLP profile ↗
ORCID search ↗
SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning
ACM MM 2024
·
Xiaoda Yang