PPaperPicks

Zehuan Yuan

14 papers at tracked venues · 12 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
  2. Goku: Flow Based Video Generative Foundation Models
  3. Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
  4. InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
  5. TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
  6. UniTok: a Unified Tokenizer for Visual Generation and Understanding
  7. EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
  8. General Object Foundation Model for Images and Videos at Scale
  9. Generative Region-Language Pretraining for Open-Ended Object Detection
  10. Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
  11. MotionMAE: Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders
  12. OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
  13. Recognize Any Regions
  14. Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction