PPaperPicks

Haizhou Li

Chinese University of Hong Kong (Shenzhen), China

77 papers at tracked venues · 40 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. CATCH: A Controllable Theme Detection Framework with Contextualized Clustering and Hierarchical Generation
  2. Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
  3. Improving Conversational Recommendation with Contextual Adaptation of External Recommenders and LLM-Based Reranking
  4. NaturalSloth: Revisiting Denial-of-Service Attacks on Large Language Models
  5. S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
  6. Towards Training-Free and Accurate ANN-to-SNN Conversion via Activation-Aware Redistribution
  7. UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-Experts
  8. Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
  9. Aligning Language Models Using Follow-up Likelihood as Reward Signal
  10. Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic Evaluation
  11. Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement Measurement
  12. Binary Event-Driven Spiking Transformer
  13. Bipolar Self-attention for Spiking Transformers
  14. Boosting Discriminability for Robust Multimodal Entity Linking with Visual Modality Missing
  15. Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
  16. ChatCRS: Incorporating External Knowledge and Goal Guidance for LLM-based Conversational Recommender Systems
  17. Decoding Listener's Identity: Person Identification from EEG Signals Using a Lightweight Spiking Transformer
  18. Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling
  19. Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural Knowledge
  20. EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph Modeling
  21. From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
  22. GTAnet: Geometry-Guided Temporal Attention for EEG-Based Sound Source Tracking in Cocktail Party Scenarios
  23. Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
  24. Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
  25. Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit Profiles
  26. Listening to the Brain: Multi-Band sEEG Auditory Reconstruction via Dynamic Spatio-Temporal Hypergraphs
  27. Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
  28. Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
  29. NeuroSpex+: Dual-Task Training of Neuro-Guided Speaker Extraction with Speech Envelope and Waveform
  30. PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
  31. Quantized Spike-driven Transformer
  32. REAL-T: Real Conversational Mixtures for Target Speaker Extraction
  33. S2NN: Sub-bit Spiking Neural Networks
  34. Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion
  35. SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
  36. SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
  37. Soundwave: Less is More for Speech-Text Alignment in LLMs
  38. SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
  39. TVC-MusicGen: Time-Varying Structure Control for Background Music Generation via Self-Supervised Training
  40. Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language Models
  41. Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
  42. TrustCLIP: Learning from Noisy Labels via Semantic Label Verification and Trust-aligned Gradient Projection
  43. UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
  44. UniTalker: Conversational Speech-Visual Synthesis
  45. Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks
  46. A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators
  47. A Non-Intrusive Approach to Assessing Dysarthria Severity: Advancing Clinical Diagnosis
  48. ASA: An Auditory Spatial Attention Dataset with Multiple Speaking Locations
  49. AceGPT, Localizing Large Language Models in Arabic
  50. Alignment at Pre-training! Towards Native Alignment for Arabic LLMs
  51. An Exploration of Length Generalization in Transformer-Based Speech Enhancement
  52. Apprenticeship-Inspired Elegance: Synergistic Knowledge Distillation Empowers Spiking Neural Networks for Efficient Single-Eye Emotion Recognition
  53. Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
  54. CMB: A Comprehensive Medical Benchmark in Chinese
  55. Does the Lombard Effect Matter in Speech Separation? Introducing the Lombard-GRID-2mix Dataset
  56. DynaThink: Fast or Slow? A Dynamic Decision-Making Framework for Large Language Models
  57. ED-sKWS: Early-Decision Spiking Neural Networks for Rapid, and Energy-Efficient Keyword Spotting
  58. Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context Modeling
  59. FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency
  60. GROOT: Generating Robust Watermark for Diffusion-Model-Based Audio Synthesis
  61. Generative Expressive Conversational Speech Synthesis
  62. How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
  63. Language Without Borders: A Dataset and Benchmark for Code-Switching Lip Reading
  64. Leveraging Graphic and Convolutional Neural Networks for Auditory Attention Detection with EEG
  65. ListenFormer: Responsive Listening Head Generation with Non-autoregressive Transformers
  66. LitE-SNN: Designing Lightweight and Efficient Spiking Neural Network through Spatial-Temporal Compressive Network Search and Joint Optimization
  67. MMAL: Multi-Modal Analytic Learning for Exemplar-Free Audio-Visual Class Incremental Tasks
  68. Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization
  69. Restoring Speaking Lips from Occlusion for Audio-Visual Speech Recognition
  70. SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
  71. SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
  72. TC-LIF: A Two-Compartment Spiking Neuron Model for Long-Term Sequential Modelling
  73. TS-Align: A Teacher-Student Collaborative Framework for Scalable Iterative Finetuning of Large Language Models
  74. UNO-DST: Leveraging Unlabelled Data in Zero-Shot Dialogue State Tracking
  75. Unveiling the Achilles' Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models
  76. WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
  77. wTIMIT2mix: A Cocktail Party Mixtures Database to Study Target Speaker Extraction for Normal and Whispered Speech