PPaperPicks

Alan L. Yuille

Johns Hopkins University, Baltimore, MD, USA

60 papers at tracked venues · 45 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. 4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
  2. 3DSRBENCH: A Comprehensive 3D Spatial Reasoning Benchmark
  3. Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency
  4. Are Pixel-Wise Metrics Reliable for Computerized Tomography Reconstruction?
  5. Autoregressive Pretraining with Mamba in Vision
  6. Baking Gaussian Splatting Into Diffusion Denoiser for Fast and Scalable Single-Stage Image-to-3D Generation and Reconstruction
  7. Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
  8. CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding
  9. Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
  10. EasyRet3D: Uncalibrated Multi-View Multi-Human 3D Reconstruction and Tracking
  11. EigenLoRAx: Recycling Adapters to Find Principal Subspaces for Resource-Efficient Adaptation and Inference
  12. FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
  13. Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution
  14. GenEx: Generating an Explorable World
  15. Learning Segmentation from Radiology Reports
  16. Mamba-Reg: Vision Mamba Also Needs Registers
  17. Medical World Model
  18. OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions
  19. PanTS: The Pancreatic Tumor Segmentation Dataset
  20. RadGPT: Constructing 3D Image-Text Tumor Datasets
  21. Scaling 3D Compositional Models for Robust Classification and Pose Estimation
  22. Scaling Laws in Patchification: An Image Is Worth 50, 176 Tokens And More
  23. Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
  24. ShapeKit
  25. Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Mutimodal Models
  26. SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
  27. SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
  28. Videoauteur: Towards Long Narrative Video Generation
  29. Vision‑Language‑Vision Auto‑Encoder: Scalable Knowledge Distillation from Diffusion Models
  30. A Bayesian Approach to OOD Robustness in Image Classification
  31. A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
  32. Benchmarking Robustness in Neural Radiance Fields
  33. Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-Modal Language Models
  34. DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
  35. De-Diffusion Makes Text a Strong Cross-Modal Interface
  36. Discovering Failure Modes of Text-guided Diffusion Models via Adversarial Search
  37. Efficient Large Multi-modal Models via Visual Context Compression
  38. Embracing Massive Medical Data
  39. From Pixel to Cancer: Cellular Automata in Computed Tomography
  40. From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
  41. Generating Images with 3D Annotations Using Diffusion Models
  42. HDR-GS: Efficient High Dynamic Range Novel View Synthesis at 1000x Speed via Gaussian Splatting
  43. HISR: Hybrid Implicit Surface Representation for Photorealistic 3D Human Reconstruction
  44. How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?
  45. IG Captioner: Information Gain Captioners Are Strong Zero-Shot Classifiers
  46. ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
  47. Masked Autoencoders are Secretly Efficient Learners
  48. NOVUM: Neural Object Volumes for Robust Object Classification
  49. NTIRE 2024 Challenge on Low Light Image Enhancement: Methods and Results
  50. Radiative Gaussian Splatting for Efficient X-Ray Novel View Synthesis
  51. Rejuvenating image-GPT as Strong Visual Representation Learners
  52. Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
  53. SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
  54. Sequential Modeling Enables Scalable Learning for Large Vision Models
  55. Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose Estimation
  56. Structure-Aware Sparse-View X-Ray 3D Reconstruction
  57. Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?
  58. Towards Generalizable Tumor Synthesis
  59. ViTamin: Designing Scalable Vision Models in the Vision-Language Era
  60. iNeMo: Incremental Neural Mesh Models for Robust Class-Incremental Learning