PPaperPicks

Dit-Yan Yeung

Hong Kong University of Science and Technology

26 papers at tracked venues · 16 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. CoherenDream: Boosting Holistic Text Coherence in 3D Generation via Multimodal Large Language Models Feedback
  2. Situated Embedding Models for Context-Aware Dense Retrieval
  3. Anyattack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models
  4. Automated Evaluation of Large Vision-Language Models on Self-Driving Corner Cases
  5. DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
  6. EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
  7. Fast and Slow Streams for Online Time Series Forecasting Without Information Leakage
  8. G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
  9. Learning 3D Persistent Embodied World Models
  10. Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
  11. The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding
  12. TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
  13. Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task
  14. DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception
  15. Eyes Closed, Safety on: Protecting Multimodal LLMs via Image-to-Text Transformation
  16. Fourier Amplitude and Correlation Loss: Beyond Using L2 Loss for Skillful Precipitation Nowcasting
  17. Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
  18. Gaussian Shell Maps for Efficient 3D Human Generation
  19. GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
  20. Implicit Concept Removal of Diffusion Models
  21. JointDreamer: Ensuring Geometry Consistency and Text Congruence in Text-to-3D Generation via Joint Score Distillation
  22. Learning High-Resolution Vector Representation from Multi-camera Images for 3D Object Detection
  23. MagicDrive: Street View Generation with Diverse 3D Geometry Control
  24. Pre-train and Refine: Towards Higher Efficiency in K-Agnostic Community Detection without Quality Degradation
  25. RoboDreamer: Learning Compositional World Models for Robot Imagination
  26. Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability