PPaperPicks

Wenhao Chai

24 papers at tracked venues · 21 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. AGLLDiff: Guiding Diffusion Models Towards Unsupervised Training-free Real-world Low-light Image Enhancement
  2. AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
    ICLR 2025 · Wenhao Chai
  3. Bringing RNNs Back to Efficient Open-Ended Video Understanding
  4. CityGen: Infinite and Controllable City Layout Generation
  5. DiffPO: Diffusion-styled Preference Optimization for Inference Time Alignment of Large Language Models
  6. Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
  7. GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning
  8. LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
  9. MonoTAKD: Teaching Assistant Knowledge Distillation for Monocular 3D Object Detection
  10. Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control is Easier than You Think
  11. PAD: Personalized Alignment of LLMs at Decoding-time
  12. PromptHaze: Prompting Real-world Dehazing via Depth Anything Model
  13. Science-T2I: Addressing Scientific Illusions in Image Synthesis
  14. ToSA: Token Merging with Spatial Awareness
  15. Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
  16. Zero-shot 3D Question Answering via Voxel-based Dynamic Token Compression
  17. Ego3DT: Tracking Every 3D Object in Ego-centric Videos
  18. LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
  19. Learning Diffusion Texture Priors for Image Restoration
  20. MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
  21. NTIRE 2024 Image Shadow Removal Challenge Report
    CVPR 2024 ·
    Florin-Alexandru Vasluianu
  22. RT-Pose: A 4D Radar Tensor-Based 3D Human Pose Estimation and Localization Benchmark
  23. See and Think: Embodied Agent in Virtual Environment
  24. UniAP: Towards Universal Animal Perception in Vision via Few-Shot Learning