PPaperPicks

Jiajun Wu

Stanford University, Stanford, CA, USA

63 papers at tracked venues · 49 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. 10 Open Challenges Steering the Future of Vision-Language-Action Models
  2. Discovering Hybrid World Representations with Co-Evolving Foundation Models
    AAAI 2026 · Jiajun Wu
  3. Birth and Death of a Rose
  4. CRAFT: Designing Creative and Functional 3D Objects
  5. Category-Agnostic Neural Object Rigging
  6. Diffusion Self-Distillation for Zero-Shot Customized Image Generation
  7. Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset
  8. Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization
  9. FluidNexus: 3D Fluid Reconstruction and Prediction from a Single Video
  10. From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
  11. Generalizable Humanoid Manipulation with 3D Diffusion Policies
  12. LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
  13. Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies
  14. Lifting Motion to the 3D World via 2D Diffusion
  15. PGC: Physics-Based Gaussian Cloth from a Single Pose
  16. Predicate Hierarchies Improve Few-Shot State Classification
  17. Range, not Independence, Drives Modularity in Biologically Inspired Representations
  18. Re-thinking Temporal Search for Long-Form Video Understanding
  19. Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals
  20. Taming generative video models for zero-shot optical flow extraction
  21. The Scene Language: Representing Scenes with Programs, Words, and Embeddings
  22. UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
  23. Understanding Complexity in VideoQA via Visual Program Generation
  24. VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
  25. Weakly-Supervised Learning of Dense Functional Correspondences
  26. Web-Scale Collection of Video Data for 4D Animal Reconstruction
  27. What Makes a Maze Look Like a Maze?
  28. Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
  29. WonderWorld: Interactive 3D Scene Generation from a Single Image
  30. Wonderplay: Dynamic 3D Scene Generation From a Single Image and Actions
  31. WorldScore: A Unified Evaluation Benchmark for World Generation
  32. X-Capture: An Open-Source Portable Device for Multi-Sensory Learning
  33. 3D Congealing: 3D-Aware Image Alignment in the Wild
  34. BEHAVIOR Vision Suite: Customizable Dataset Generation via Simulation
  35. CityPulse: Fine-Grained Assessment of Urban Change with Street View Time Series
  36. Controllable Human-Object Interaction Synthesis
  37. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
  38. Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
  39. FactorSim: Generative Simulation via Factorized Representation
  40. Hearing Anything Anywhere
  41. Holodeck: Language Guided Generation of 3D Embodied AI Environments
  42. HourVideo: 1-Hour Video-Language Understanding
  43. IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
  44. Language-Informed Visual Concept Learning
  45. Learning Planning Abstractions from Language
  46. Learning the 3D Fauna of the Web
  47. Learning to Design 3D Printable Adaptations on Everyday Objects for Robot Manipulation
  48. MARPLE: A Benchmark for Long-Horizon Inference
  49. Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
  50. Neural Polynomial Gabor Fields for Macro Motion Analysis
  51. Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration
  52. Patched Denoising Diffusion Models For High-Resolution Image Synthesis
  53. PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation
  54. Physically Grounded Vision-Language Models for Robotic Manipulation
  55. Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos
  56. Reconstruction and Simulation of Elastic Objects with Spring-Mass 3D Gaussians
  57. RoboPack: Learning Tactile-Informed Dynamics Models for Dense Packing
  58. SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing
  59. Streaming Detection of Queried Event Start
  60. Tripod: Three Complementary Inductive Biases for Disentangled Representation Learning
  61. ULIP-2: Towards Scalable Multimodal Pre-Training for 3D Understanding
  62. WonderJourney: Going from Anywhere to Everywhere
  63. ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image