PPaperPicks

Jing Shi

14 papers at tracked venues · 12 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Plot'n Polish: Zero-Shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
  2. AV-DiT: Taming Image Diffusion Transformers for Efficient Joint Audio and Video Generation
  3. DiffTell: A High-Quality Dataset for Describing Image Manipulation Changes
  4. FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity
  5. GUI Agents: A Survey
  6. Improving Large Vision and Language Models by Learning from a Panel of Peers
  7. MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities
  8. The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers
  9. Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
  10. Visual Persona: Foundation Model for Full-Body Human Customization
  11. Yo'Chameleon: Personalized Vision and Language Generation
  12. Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
  13. FineMatch: Aspect-Based Fine-Grained Image and Text Mismatch Detection and Correction
  14. VIXEN: Visual Text Comparison Network for Image Difference Captioning