PPaperPicks

Zhihong Zhu

Tencent Jarvis Lab, Shenzhen, China

53 papers at tracked venues · 28 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Beyond Surface Features: Advancing Medical Vision-Language Alignment via Dynamic Evidence-Guided Preference Optimization
  2. CMID: Towards Medical Visual Question Answering via Contrastive Mutual Information Decoding
    AAAI 2026 · Zhihong Zhu
  3. MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models
  4. Multimodal Dual-Path Decoding for Medical Report Generation
  5. S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm Detection
    AAAI 2026 · Zhihong Zhu
  6. A Survey on Foundation Language Models for Single-cell Biology
  7. A Survey on Multi-modal Intent Recognition: Recent Advances and New Frontiers
    EMNLP 2025 · Zhihong Zhu
  8. CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLM
  9. Can We Trust AI Doctors? A Survey of Medical Hallucination in Large Language and Large Vision-Language Models
    ACL 2025 · Zhihong Zhu
  10. CellVerse: Do Large Language Models Really Understand Cell Biology?
  11. D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
  12. DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
  13. HTML: Hierarchical Topology Multi-task Learning for Semantic Parsing in Knowledge Base Question Answering
  14. Harnessing Large Language Models for Knowledge Graph Question Answering via Adaptive Multi-Aspect Retrieval-Augmentation
  15. MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
  16. PgM: Partitioner Guided Modal Learning Framework
  17. RTE-GMoE: A Model-agnostic Approach for Relation Triplet Extraction via Graph-based Mixture-of-Expert Mutual Learning
  18. UniCoTT: A Unified Framework for Structural Chain-of-Thought Distillation
  19. VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification
  20. Aligner²: Enhancing Joint Multiple Intent Detection and Slot Filling via Adjustive and Forced Cross-Task Alignment
    AAAI 2024 · Zhihong Zhu
  21. Aspects are Anchors: Towards Multimodal Aspect-based Sentiment Analysis via Aspect-driven Alignment and Refinement
  22. Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
  23. AutoPRM: Automating Procedural Supervision for Multi-Step Reasoning via Controllable Question Decomposition
  24. Code-Switching Can be Better Aligners: Advancing Cross-Lingual SLU through Representation-Level and Prediction-Level Alignment
    ACL 2024 · Zhihong Zhu
  25. Cyclical Contrastive Learning Based on Geodesic for Zero-shot Cross-lingual Spoken Language Understanding
  26. DGLF: A Dual Graph-based Learning Framework for Multi-modal Sarcasm Detection
    EMNLP 2024 · Zhihong Zhu
  27. Dance with Labels: Dual-Heterogeneous Label Graph Interaction for Multi-intent Spoken Language Understanding
    WSDM 2024 · Zhihong Zhu
  28. DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
  29. Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent Detection
  30. Editing Factual Knowledge and Explanatory Ability of Medical Large Language Models
  31. Exploiting Auxiliary Caption for Video Grounding
  32. GPA: Global and Prototype Alignment for Audio-Text Retrieval
  33. Game on Tree: Visual Hallucination Mitigation via Coarse-to-Fine View Tree and Game Theory
  34. InMu-Net: Advancing Multi-modal Intent Detection via Information Bottleneck and Multi-sensory Processing
    ACM MM 2024 · Zhihong Zhu
  35. KDProR: A Knowledge-Decoupling Probabilistic Framework for Video-Text Retrieval
  36. LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
  37. Learning to Match Representations is Better for End-to-End Task-Oriented Dialog System
  38. MOAT: Graph Prompting for 3D Molecular Graphs
  39. MedJourney: Benchmark and Evaluation of Large Language Models over Patient Clinical Journey
  40. Mitigating Hallucinations of Large Language Models in Medical Information Extraction via Contrastive Decoding
  41. MoBA: Mixture of Bi-directional Adapter for Multi-modal Sarcasm Detection
  42. MoE-SLU: Towards ASR-Robust Spoken Language Understanding via Mixture-of-Experts
  43. Multivariate Cooperative Game for Image-Report Pairs: Hierarchical Semantic Alignment for Medical Report Generation
    MICCAI 2024 · Zhihong Zhu
  44. PIXEL: Prompt-based Zero-shot Hashing via Visual and Textual Semantic Alignment
  45. Relevance Is a Guiding Light: Relevance-aware Adaptive Learning for End-to-end Task-oriented Dialogue System
  46. SaLa: Scenario-aware Label Graph Interaction for Multi-intent Spoken Language Understanding
    CIKM 2024 · Zhihong Zhu
  47. TFCD: Towards Multi-modal Sarcasm Detection via Training-Free Counterfactual Debiasing
    IJCAI 2024 · Zhihong Zhu
  48. Textual Inversion and Self-supervised Refinement for Radiology Report Generation
  49. Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal Transport
  50. Towards Multimodal-augmented Pre-trained Language Models via Self-balanced Expectation-Maximization Iteration
  51. UniMEEC: Towards Unified Multimodal Emotion Recognition and Emotion Cause
  52. What are the Generator Preferences for End-to-end Task-Oriented Dialog System?
  53. XMeCap: Meme Caption Generation with Sub-Image Adaptability