PPaperPicks

Seong-Whan Lee

Korea University, Seoul, South Korea

22 papers at tracked venues · 15 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
  2. ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment
  3. Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
  4. DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
  5. DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
  6. EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
  7. FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Control
  8. MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
  9. PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
  10. PoseAnchor: Robust Root Position Estimation for 3D Human Pose Estimation
  11. ProPose: Probabilistic 3D Human Pose Estimation with Instance-Level Distribution and Normalizing Flow
  12. Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
  13. Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency Partition
  14. Towards Generalizable 3D Human Pose Estimation via Ensembles on Flat Loss Landscapes
  15. VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
  16. XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answering
  17. DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice Conversion
  18. EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
  19. TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression
  20. Text-Infused Attention and Foreground-Aware Modeling for Zero-Shot Temporal Action Detection
  21. Toward Approaches to Scalability in 3D Human Pose Estimation
  22. Unknown-Aware Graph Regularization for Robust Semi-supervised Learning from Uncurated Data