P
PaperPicks
Conferences
Shiliang Zhang
30 papers at tracked venues · 23 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
AAAI
×7
CVPR
×6
InterSpeech
×6
ACL
×3
ACM MM
×3
ICML
×2
EMNLP
×1
ICCV
×1
NeurIPS
×1
Frequent coauthors
Zehong Ma
DBLP profile ↗
ORCID search ↗
×3
Dongkai Wang
DBLP profile ↗
ORCID search ↗
×3
Zhanzhou Feng
DBLP profile ↗
ORCID search ↗
×2
Guanrou Yang
DBLP profile ↗
ORCID search ↗
×2
Junwei Zhao
DBLP profile ↗
ORCID search ↗
×2
Ziyang Ma
DBLP profile ↗
ORCID search ↗
×2
Shiyu Xuan
DBLP profile ↗
ORCID search ↗
×2
Xiao Wang
DBLP profile ↗
ORCID search ↗
×1
Changfeng Gao
DBLP profile ↗
ORCID search ↗
×1
Mingyu Cui
DBLP profile ↗
ORCID search ↗
×1
Haoyu Wang
DBLP profile ↗
ORCID search ↗
×1
Ruihan Xu
DBLP profile ↗
ORCID search ↗
×1
Papers
SCAN: Self-Calibrated AutoregressioN for High-Quality Visual Generation
AAAI 2026
·
Zhanzhou Feng
DBLP profile ↗
ORCID search ↗
When Person Re-Identification Meets Event Camera: A Benchmark Dataset and an Attribute-Guided Re-Identification Framework
AAAI 2026
·
Xiao Wang
DBLP profile ↗
ORCID search ↗
Differentiable Reward Optimization for LLM based TTS system
InterSpeech 2025
·
Changfeng Gao
DBLP profile ↗
ORCID search ↗
Efficient Multi-modal Long Context Learning for Training-free Adaptation
ICML 2025
·
Zehong Ma
DBLP profile ↗
ORCID search ↗
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
ACM MM 2025
·
Guanrou Yang
DBLP profile ↗
ORCID search ↗
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
InterSpeech 2025
·
Mingyu Cui
DBLP profile ↗
ORCID search ↗
Generalizable Object Keypoint Localization from Generative Priors
CVPR 2025
·
Dongkai Wang
DBLP profile ↗
ORCID search ↗
HDCFN: Haze Distribution-aware Cross-modal Fusion Network for Infrared-guided Dense Haze Removal in UAVs
ACM MM 2025
·
Junwei Zhao
DBLP profile ↗
ORCID search ↗
MV-VTON: Multi-View Virtual Try-On with Diffusion Models
AAAI 2025
·
Haoyu Wang
DBLP profile ↗
ORCID search ↗
MagCache: Fast Video Generation with Magnitude-Aware Cache
NeurIPS 2025
·
Zehong Ma
DBLP profile ↗
ORCID search ↗
NN-Former: Rethinking Graph Structure in Neural Architecture Representation
CVPR 2025
·
Ruihan Xu
DBLP profile ↗
ORCID search ↗
OmniAudio: Generating Spatial Audio from 360-Degree Video
ICML 2025
·
Huadai Liu
DBLP profile ↗
ORCID search ↗
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
ACL 2025
·
Qinglin Zhang
DBLP profile ↗
ORCID search ↗
Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration
AAAI 2025
·
Ziyang Ma
DBLP profile ↗
ORCID search ↗
UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
ACL 2025
·
Yidi Jiang
DBLP profile ↗
ORCID search ↗
UniSpeaker: A Unified Approach for Multimodality-driven Speaker Generation
EMNLP 2025
·
Zhengyan Sheng
DBLP profile ↗
ORCID search ↗
Unified Visual Generation via Next-Set Prediction in Continuous Domain
ICCV 2025
·
Zhanzhou Feng
DBLP profile ↗
ORCID search ↗
CoTuning: A Large-Small Model Collaborating Distillation Framework for Better Model Generalization
ACM MM 2024
·
Zimo Liu
DBLP profile ↗
ORCID search ↗
Decoupled Contrastive Learning for Long-Tailed Recognition
AAAI 2024
·
Shiyu Xuan
DBLP profile ↗
ORCID search ↗
Decoupled Optimisation for Long-Tailed Visual Recognition
AAAI 2024
·
Cong Cong
DBLP profile ↗
ORCID search ↗
ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
InterSpeech 2024
·
Yafeng Chen
DBLP profile ↗
ORCID search ↗
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
InterSpeech 2024
·
Peng Wang
DBLP profile ↗
ORCID search ↗
LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model
CVPR 2024
·
Dongkai Wang
DBLP profile ↗
ORCID search ↗
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
InterSpeech 2024
·
Guanrou Yang
DBLP profile ↗
ORCID search ↗
OVMR: Open-Vocabulary Recognition with Multi-Modal References
CVPR 2024
·
Zehong Ma
DBLP profile ↗
ORCID search ↗
Personality-memory Gated Adaptation: An Efficient Speaker Adaptation for Personalized End-to-end Automatic Speech Recognition
InterSpeech 2024
·
Yue Gu
DBLP profile ↗
ORCID search ↗
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
CVPR 2024
·
Shiyu Xuan
DBLP profile ↗
ORCID search ↗
Recognizing Ultra-High-Speed Moving Objects with Bio-Inspired Spike Camera
AAAI 2024
·
Junwei Zhao
DBLP profile ↗
ORCID search ↗
Spatial-Aware Regression for Keypoint Localization
CVPR 2024
·
Dongkai Wang
DBLP profile ↗
ORCID search ↗
emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
ACL 2024
·
Ziyang Ma
DBLP profile ↗
ORCID search ↗