PPaperPicks

Xiaoda Yang

22 papers at tracked venues · 15 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. SpatialLogic-Bench: A Diagnostic Benchmark for Task-Oriented Spatiotemporal Reasoning
    AAAI 2026 · Xiaoda Yang
  2. Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
  3. VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
  4. BrainLoc: Brain Signal-Based Object Detection with Multi-modal Alignment
  5. CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling
  6. Choose Your Expert: Uncertainty-Guided Expert Selection for Continual Deepfake Detection
  7. Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
  8. EAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic Integration
  9. EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
  10. GTA: Towards Generative Text-To-Audio Retrieval via Multi-Scale Tokenizer
  11. MelRe: Vision-Based Mel-Spectrogram Restoration
  12. Multimodal Conditional Retrieval with High Controllability
    SIGKDD 2025 · Xiaoda Yang
  13. PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue
  14. Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
  15. Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
  16. Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection
  17. Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
  18. VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?
  19. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
  20. AudioVSR: Enhancing Video Speech Recognition with Audio Data
    EMNLP 2024 · Xiaoda Yang
  21. Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented Prompts
  22. SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning
    ACM MM 2024 · Xiaoda Yang