PPaperPicks

Xuxin Cheng

41 papers at tracked venues · 24 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
  2. Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
  3. MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
  4. SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
  5. When 20 Agents Fail to Sort: The Distributed Sorting Benchmark for Scalable Multi-Agent Systems
  6. CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model
  7. DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
  8. EXCGEC: A Benchmark for Edit-Wise Explainable Chinese Grammatical Error Correction
  9. Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
  10. Mobile-TeleVision: Predictive Motion Priors for Humanoid Whole-Body Control
  11. UniCoTT: A Unified Framework for Structural Chain-of-Thought Distillation
  12. Aligner²: Enhancing Joint Multiple Intent Detection and Slot Filling via Adjustive and Forced Cross-Task Alignment
  13. Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
  14. Code-Switching Can be Better Aligners: Advancing Cross-Lingual SLU through Representation-Level and Prediction-Level Alignment
  15. Cyclical Contrastive Learning Based on Geodesic for Zero-shot Cross-lingual Spoken Language Understanding
    ACL 2024 · Xuxin Cheng
  16. Dance with Labels: Dual-Heterogeneous Label Graph Interaction for Multi-intent Spoken Language Understanding
  17. DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
  18. Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
  19. Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation
  20. Exploiting Auxiliary Caption for Video Grounding
  21. Expressive Whole-Body Control for Humanoid Robots
    RSS 2024 · Xuxin Cheng
  22. Extreme Parkour with Legged Robots
    ICRA 2024 · Xuxin Cheng
  23. FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
  24. Generating More Audios for End-to-End Spoken Language Understanding
    IJCAI 2024 · Xuxin Cheng
  25. InMu-Net: Advancing Multi-modal Intent Detection via Information Bottleneck and Multi-sensory Processing
  26. KDProR: A Knowledge-Decoupling Probabilistic Framework for Video-Text Retrieval
  27. Learning to Match Representations is Better for End-to-End Task-Oriented Dialog System
  28. MaCSC: Towards Multimodal-augmented Pre-trained Language Models via Conceptual Prototypes and Self-balancing Calibration
  29. MoE-SLU: Towards ASR-Robust Spoken Language Understanding via Mixture-of-Experts
    ACL 2024 · Xuxin Cheng
  30. Multivariate Cooperative Game for Image-Report Pairs: Hierarchical Semantic Alignment for Medical Report Generation
  31. PCAD: Towards ASR-Robust Spoken Language Understanding via Prototype Calibration and Asymmetric Decoupling
  32. PolyVoice: Language Models for Speech to Speech Translation
  33. RAG-HAT: A Hallucination-Aware Tuning Pipeline for LLM in Retrieval-Augmented Generation
  34. Retrieval is Accurate Generation
  35. SaLa: Scenario-aware Label Graph Interaction for Multi-intent Spoken Language Understanding
  36. Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup
    ACL 2024 · Xuxin Cheng
  37. Towards Explainable Joint Models via Information Theory for Multiple Intent Detection and Slot Filling
  38. Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal Transport
    AAAI 2024 · Xuxin Cheng
  39. Towards Multimodal-augmented Pre-trained Language Models via Self-balanced Expectation-Maximization Iteration
  40. Uncertainty-Aware Sign Language Video Retrieval with Probability Distribution Modeling
  41. What are the Generator Preferences for End-to-end Task-Oriented Dialog System?