PPaperPicks

Jianhua Han

21 papers at tracked venues · 15 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. DisCo: Discovering Common Affordance from Large Models for Actionable Part Perception
  2. EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
  3. G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
  4. HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
  5. ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
  6. SeePhys: Does Seeing Help Thinking? - Benchmarking Vision-Based Physics Reasoning
  7. Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
  8. Any-Size-Diffusion: Toward Efficient Text-Driven Synthesis for Any-Size HD Images
  9. CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation
  10. DetCLIPv3: Towards Versatile Generative Open-Vocabulary Object Detection
  11. Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
  12. Holistic Autonomous Driving Understanding by Bird'View Injected Multi-Modal Large Models
  13. HumanRefiner: Benchmarking Abnormal Human Generation and Refining with Coarse-to-Fine Pose-Reversible Guidance
  14. Implicit Concept Removal of Diffusion Models
  15. Ins-DetCLIP: Aligning Detection Model to Follow Human-Language Instruction
  16. LayerDiff: Exploring Text-Guided Multi-layered Composable Image Synthesis via Layer-Collaborative Diffusion Model
  17. PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
  18. Reason2Drive: Towards Interpretable and Chain-Based Reasoning for Autonomous Driving
  19. SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
  20. UNIT: Unifying Image and Text Recognition in One Vision Encoder
  21. VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation