PPaperPicks

Xintao Wang

Tencent AI Lab., Tencent PCG, Applied Research Center (ARC), Shenzhen, China

39 papers at tracked venues · 33 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation
  2. Anti-Diffusion: Preventing Abuse of Modifications of Diffusion-Based Models
  3. CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
  4. Flow-GRPO: Training Flow Matching Models via Online RL
  5. FullDiT: Video Generative Foundation Models with Multimodal Control via Full Attention
  6. GameFactorly: Creating New Games with Generative Interactive Videos
  7. Image Conductor: Precision Control for Interactive Video Synthesis
  8. Improving Video Generation with Human Feedback
  9. MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
  10. PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
  11. Recammaster: Camera-Controlled Generative Rendering From a Single Video
  12. SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
  13. BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion
  14. CustomNet: Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models
  15. DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing
  16. DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models
  17. DreamDiffusion: High-Quality EEG-to-Image Generation with Temporal Masked Signal Modeling and CLIP Alignment
  18. DynamiCrafter: Animating Open-Domain Images with Video Diffusion Priors
  19. EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
  20. Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free Videos
  21. FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling
  22. MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
  23. Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation
  24. Making LLaMA SEE and Draw with SEED Tokenizer
  25. MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
  26. PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
  27. ReVideo: Remake a Video with Motion and Content Control
  28. Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis
  29. ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion Models
  30. Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild
  31. Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
  32. SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models
  33. SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model
  34. Storytelling Video Generation with Retrieval Augmentation and Character Consistency
  35. T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion Models
  36. Unifying Image Processing as Visual Prompting Question Answering
  37. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
  38. VideoTetris: Towards Compositional Text-to-Video Generation
  39. X- Adapter: Universal Compatibility of Plugins for Upgraded Diffusion Model