PPaperPicks

Ying Shan

82 papers at tracked venues · 68 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
  2. MMhops-R1: Multimodal Multi-hop Reasoning
  3. AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
  4. CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
  5. DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creation
  6. DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
  7. DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation
  8. DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
  9. Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
  10. FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction
  11. GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers
  12. Geometrycrafter: Consistent Geometry Estimation for Open-World Videos With Diffusion Priors
  13. HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
  14. Image Conductor: Precision Control for Interactive Video Synthesis
  15. LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
  16. Mamba-3VL: Taming State Space Model for 3D Vision Language Learning
  17. Mani-GS: Gaussian Splatting Manipulation with Triangular Mesh
  18. MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
  19. Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion
  20. Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
  21. NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images
  22. Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
  23. SEED-Story: Multimodal Long Story Generation with Large Language Model
  24. Scalable Image Tokenization with Index Backpropagation Quantization
  25. Taming Rectified Flow for Inversion and Editing
  26. TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
  27. UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
  28. VisionMath: Vision-Form Mathematical Problem-Solving
  29. A Pre-convolved Representation for Plug-and-Play Neural Illumination Fields
  30. AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
  31. BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
  32. BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion
  33. CV-VAE: A Compatible Video VAE for Latent Generative Video Models
  34. ConTex-Human: Free-View Rendering of Human from a Single Image with Texture-Consistent Synthesis
  35. CustomNet: Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models
  36. DMiT: Deformable Mipmapped Tri-Plane Representation for Dynamic Scenes
  37. DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing
  38. DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models
  39. DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion Models
  40. DreamDiffusion: High-Quality EEG-to-Image Generation with Temporal Masked Signal Modeling and CLIP Alignment
  41. DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing
  42. DynamiCrafter: Animating Open-Domain Images with Video Diffusion Priors
  43. E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
  44. EA-VTR: Event-Aware Video-Text Retrieval
  45. EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
  46. FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling
  47. GS-IR: 3D Gaussian Splatting for Inverse Rendering
  48. HiFi-123: Towards High-Fidelity One Image to 3D Content Generation
  49. How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?
  50. HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting
  51. HumanRef: Single Image to 3D Human Generation via Reference-Guided Diffusion
  52. LLaMA Pro: Progressive LLaMA with Block Expansion
  53. Low-Rank Approximation for Sparse Attention in Multi-Modal LLMs
  54. MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
  55. Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation
  56. Making LLaMA SEE and Draw with SEED Tokenizer
  57. MambaTree: Tree Topology is All You Need in State Space Model
  58. MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
  59. Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
  60. Noise Calibration: Plug-and-Play Content-Preserving Video Enhancement Using Pre-trained Video Diffusion Models
  61. PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
  62. Programmable Motion Generation for Open-Set Motion Control Tasks
  63. ReVideo: Remake a Video with Motion and Content Control
  64. RecDCL: Dual Contrastive Learning for Recommendation
  65. Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis
  66. SC-NeuS: Consistent Neural Surface Reconstruction from Sparse and Noisy Views
  67. SEED-Bench: Benchmarking Multimodal Large Language Models
  68. ST-LLM: Large Language Models Are Effective Temporal Learners
  69. ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion Models
  70. SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models
  71. Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse Views
  72. SparseGNV: Generating Novel Views of Indoor Scenes with Sparse RGB-D Images
  73. SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model
  74. Storytelling Video Generation with Retrieval Augmentation and Character Consistency
  75. SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
  76. T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion Models
  77. TapMo: Shape-aware Motion Generation of Skeleton-free Characters
  78. Texture-GS: Disentangling the Geometry and Texture for 3D Gaussian Splatting Editing
  79. UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
  80. VIT-LENS: Towards Omni-modal Representations
  81. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
  82. YOLO-World: Real-Time Open-Vocabulary Object Detection