PPaperPicks

Xiaoming Wei

12 papers at tracked venues · 12 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal Diffusion
  2. ARIG: Autoregressive Interactive Head Generation for Real-Time Conversations
  3. Denoising with a Joint-Embedding Predictive Architecture
  4. LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
  5. Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
  6. Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models
  7. Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
  8. Animating General Image with Large Visual Motion Model
  9. BEM: Balanced and Entropy-Based Mix for Long-Tailed Semi-Supervised Learning
  10. Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
  11. ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting
  12. Real3D: The Curious Case of Neural Scene Degeneration