P
PaperPicks
Conferences
Ryo Masumura
14 papers at tracked venues · 4 at CORE A* · active 2024–2026
DBLP profile ↗
ORCID search ↗
Venues
InterSpeech
×9
AAAI
×3
ICCV
×1
WACV
×1
Frequent coauthors
Naoki Makishima
DBLP profile ↗
ORCID search ↗
×3
Satoshi Suzuki
DBLP profile ↗
ORCID search ↗
×2
Kazutoshi Shinoda
DBLP profile ↗
ORCID search ↗
×2
Haris Gulzar
DBLP profile ↗
ORCID search ↗
×1
Taiga Yamane
DBLP profile ↗
ORCID search ↗
×1
Takafumi Moriya
DBLP profile ↗
ORCID search ↗
×1
Atsushi Ando
DBLP profile ↗
ORCID search ↗
×1
Keita Suzuki
DBLP profile ↗
ORCID search ↗
×1
Papers
Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models
AAAI 2026
·
Satoshi Suzuki
DBLP profile ↗
ORCID search ↗
Distribution Highlighted Reference-based Label Distribution Learning for Facial Age Estimation
WACV 2026
·
Satoshi Suzuki
DBLP profile ↗
ORCID search ↗
Leveraging LLMs for Written to Spoken Style Data Transformation to Enhance Spoken Dialog State Tracking
InterSpeech 2025
·
Haris Gulzar
DBLP profile ↗
ORCID search ↗
MVTrajecter: Multi-View Pedestrian Tracking With Trajectory Motion Cost and Trajectory Appearance Cost
ICCV 2025
·
Taiga Yamane
DBLP profile ↗
ORCID search ↗
Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level Scores
AAAI 2025
·
Ryo Masumura
SOMSRED-SVC: Sequential Output Modeling with Speaker Vector Constraints for Joint Multi-Talker Overlapped ASR and Speaker Diarization
InterSpeech 2025
·
Naoki Makishima
DBLP profile ↗
ORCID search ↗
ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind
AAAI 2025
·
Kazutoshi Shinoda
DBLP profile ↗
ORCID search ↗
Unified Audio-Visual Modeling for Recognizing Which Face Spoke When and What in Multi-Talker Overlapped Speech and Video
InterSpeech 2025
·
Naoki Makishima
DBLP profile ↗
ORCID search ↗
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
InterSpeech 2024
·
Takafumi Moriya
DBLP profile ↗
ORCID search ↗
Factor-Conditioned Speaking-Style Captioning
InterSpeech 2024
·
Atsushi Ando
DBLP profile ↗
ORCID search ↗
Learning from Multiple Annotator Biased Labels in Multimodal Conversation
InterSpeech 2024
·
Kazutoshi Shinoda
DBLP profile ↗
ORCID search ↗
Participant-Pair-Wise Bottleneck Transformer for Engagement Estimation from Video Conversation
InterSpeech 2024
·
Keita Suzuki
DBLP profile ↗
ORCID search ↗
SOMSRED: Sequential Output Modeling for Joint Multi-talker Overlapped Speech Recognition and Speaker Diarization
InterSpeech 2024
·
Naoki Makishima
DBLP profile ↗
ORCID search ↗
Unified Multi-Talker ASR with and without Target-speaker Enrollment
InterSpeech 2024
·
Ryo Masumura