PPaperPicks

Zuxuan Wu

43 papers at tracked venues · 37 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. DriveSuprim: Towards Precise Trajectory Selection for End-to-End Planning
  2. Human2Robot: Learning Robot Actions from Paired Human-Robot Videos
  3. Achieving More with Less: Additive Prompt Tuning for Rehearsal-Free Class-Incremental Learning
  4. AdaDiff: Adaptive Step Selection for Fast Diffusion Models
  5. Adaptive Retention & Correction: Test-Time Training for Continual Learning
  6. AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments
  7. Aid: Adapting Image2video Diffusion Models for Instruction-Guided Video Prediction
  8. BlockDance: Reuse Structurally Similar Spatio-Temporal Features to Accelerate Diffusion Transformers
  9. Comprehensive Multi-Modal Prototypes Are Simple and Effective Classifiers for Vast-Vocabulary Object Detection
  10. CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
  11. EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation
  12. FNIN: A Fourier Neural Operator-based Numerical Integration Network for Surface-from-gradients
  13. FOCUS: Towards Universal Foreground Segmentation
  14. ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation Detection
  15. Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
  16. INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
  17. MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
  18. MotionFollower: Editing Video Motion via Score-Guided Diffusion
  19. OmniGen-AR: AutoRegressive Any-to-Image Generation
  20. ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning
  21. REDUCIO! Generating 1K Video Within 16 Seconds Using Extremely Compressed Motion Latents
  22. Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
  23. Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control
  24. StableAnimator: High-Quality Identity-Preserving Human Image Animation
  25. UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
  26. VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
  27. Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms
  28. BEVNeXt: Reviving Dense BEV Frameworks for 3D Object Detection
  29. DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
  30. DreamMesh: Jointly Manipulating and Texturing Triangle Meshes for Text-to-3D Generation
  31. Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
  32. GenRec: Unifying Video Generation and Recognition with Diffusion Models
  33. Learning to Rank Patches for Unbiased Image Redundancy Reduction
  34. MagDiff: Multi-alignment Diffusion for High-Fidelity Video Generation and Editing
  35. ModelLock: Locking Your Model With a Spell
  36. MotionEditor: Editing Video Motion via Content-Aware Diffusion
  37. OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
  38. OmniViD: A Generative Framework for Universal Video Understanding
  39. PromptFusion: Decoupling Stability and Plasticity for Continual Learning
  40. SegIC: Unleashing the Emergent Correspondence for In-Context Segmentation
  41. SimDA: Simple Diffusion Adapter for Efficient Video Generation
  42. Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding
  43. Zero-shot High-fidelity and Pose-controllable Character Animation