PPaperPicks

Ming Hu

31 papers at tracked venues · 21 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. CNText2Sign and CNSign: Unified Chinese Sign Language Datasets for Bidirectional Accessibility
  2. GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI
  3. Hierarchical Prompt Contrastive Learning for Weakly Supervised Histopathology Segmentation
  4. S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything Without Supervision
  5. SPIDE: Serial and Parallel Intertwined Speculative Decoding
  6. TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification
  7. DONIS: Importance Sampling for Training Physics-Informed DeepONet
  8. Decoding Causal Structure: End-to-End Mediation Pathways Inference
  9. Derm1M: A Million-Scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology
  10. Genesis: A Large-Scale Benchmark for Multimodal Large Language Model in Emotional Causality Analysis
  11. Local Masked Reconstruction for Efficient Self-Supervised Learning on High-Resolution Images
  12. MAKE: Multi-Aspect Knowledge-Enhanced Vision-Language Pretraining for Zero-Shot Dermatological Assessment
  13. MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation
  14. MSWAL: 3D Multi-class Segmentation of Whole Abdominal Lesions Dataset
  15. Neighbor Does Matter: Density-Aware Contrastive Learning for Medical Semi-supervised Segmentation
  16. One Arrow, Two Hawks: Sharpness-aware Minimization for Federated Learning via Global Model Trajectory
  17. OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
    ICCV 2025 · Ming Hu
  18. Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model
  19. Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level Guidance
  20. RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions
  21. Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation
  22. Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
  23. SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding
  24. Star with Bilinear Mapping
  25. Temporal Model-Based Federated Active Medical Image Classification
  26. Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery
    NeurIPS 2025 · Ming Hu
  27. Towards Realistic Semi-supervised Medical Image Classification
  28. Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection
  29. UniViT: Unifying Image and Video Understanding in One Vision Encoder
  30. Generalizing to Unseen Domains in Diabetic Retinopathy with Disentangled Representations
  31. OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding
    ECCV 2024 · Ming Hu