PPaperPicks

Pheng-Ann Heng

Chinese University of Hong Kong, Hong Kong

51 papers at tracked venues · 27 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation
  2. SurgPub-Video: A Comprehensive Surgical Video Framework for Enhanced Surgical Intelligence in Vision-Language Model
  3. COS3D: Collaborative Open-Vocabulary 3D Segmentation
  4. CellVerse: Do Large Language Models Really Understand Cell Biology?
  5. ChemMiner: A Large Language Model Agent System for Chemical Literature Data Mining
  6. ClipGS: Clippable Gaussian Splatting for Interactive Cinematic Visualization of Volumetric Medical Data
  7. Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
  8. DiTAC: Discrete Teamwork Abstraction for Ad Hoc Collaboration
  9. Dual Ensembled Multiagent Q-Learning with Hypernet Regularizer
  10. EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights
  11. Fast Image Super-Resolution via Consistency Rectified Flow
  12. Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning
  13. GLID$2$E: A Gradient-Free Lightweight Fine-tune Approach for Discrete Biological Sequence Design
  14. MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
  15. MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models
  16. Medical Large Vision Language Models with Multi-image Visual Ability
  17. Protein Inverse Folding From Structure Feedback
  18. Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splatting
  19. SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
  20. SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
  21. Sequence-Independent Continual Test-Time Adaptation with Mixture of Incremental Experts for Cross-Domain Segmentation
  22. Surgical Workflow Recognition and Blocking Effectiveness Detection in Laparoscopic Liver Resection with Pringle Maneuver
  23. T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
  24. Topology-Constrained Learning for Efficient Laparoscopic Liver Landmark Detection
  25. UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation
  26. Video Grounded Conversation Generation for Reference Surgical Instrument Segmentation
  27. What We Miss Matters: Learning from the Overlooked in Point Cloud Transformers
  28. ANEDL: Adaptive Negative Evidential Deep Learning for Open-Set Semi-supervised Learning
  29. Coarse-to-Fine Latent Diffusion Model for Glaucoma Forecast on Sequential Fundus Images
  30. Comprehensive Generative Replay for Task-Incremental Segmentation with Concurrent Appearance and Semantic Forgetting
  31. Cross Prompting Consistency with Segment Anything Model for Semi-supervised Medical Image Segmentation
  32. DR-Label: Label Deconstruction and Reconstruction of GNN Models for Catalysis Systems
  33. Depth-Driven Geometric Prompt Learning for Laparoscopic Liver Landmark Detection
  34. Distribution-Aware Calibration for Object Detection with Noisy Bounding Boxes
  35. Epicardium Prompt-Guided Real-Time Cardiac Ultrasound Frame-to-Volume Registration
  36. FM-OSD: Foundation Model-Enabled One-Shot Detection of Anatomical Landmarks
  37. LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic Surgery
  38. LoRAExit: Empowering Dynamic Modulation of LLMs in Resource-limited Settings using Low-rank Adapters
  39. Memory-Efficient Prompt Tuning for Incremental Histopathology Classification
  40. Neural P3M: A Long-Range Interaction Modeling Enhancer for Geometric GNNs
  41. Noise Level Adaptive Diffusion Model for Robust Reconstruction of Accelerated MRI
  42. PCF-Lift: Panoptic Lifting by Probabilistic Contrastive Fusion
  43. PointPatchMix: Point Cloud Mixing with Patch Scoring
  44. Sample-Efficient Multiagent Reinforcement Learning with Reset Replay
  45. SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View Adaptation
  46. SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
  47. Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models
  48. Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement Learning
  49. Tri-Modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms
  50. Unveiling the Generalization Power of Fine-Tuned Large Language Models
  51. Weakly-Supervised Medical Image Segmentation with Gaze Annotations