PPaperPicks

Dimitris N. Metaxas

Rutgers University, Division of Computer and Information Sciences, Piscataway, NJ, USA

42 papers at tracked venues · 31 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
  2. DICE: Discrete Inversion Enabling Controllable Editing for Masked Generative Models
  3. Individual Turing Test: A Case Study of LLM-based Simulation Using Longitudinal Personal Data
  4. Large Sign Language Models: Toward 3D American Sign Language Translation
  5. Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
  6. Snapmoji: Instant Generation of Animatable Dual-Stylized Avatars
  7. Stable Signer: Hierarchical Sign Language Generative Model
  8. APEER : Automatic Prompt Engineering Enhances Large Language Model Reranking
  9. Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction
  10. AutoEdit: Automatic Hyperparameter Tuning for Image Editing
  11. Continuous Spatio-Temporal Memory Networks for 4D Cardiac Cine MRI Segmentation
  12. FlowChef: Steering of Rectified Flow Models for Controlled Generations
  13. Implicit In-context Learning
  14. Improved Training Technique for Latent Consistency Models
  15. LUCAS: Layered Universal Codec Avatars
  16. LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
  17. MLLM-as-a-Judge for Image Safety without Human Labeling
  18. RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
  19. RankFlow: A Multi-Role Collaborative Reranking Workflow Utilizing Large Language Models
  20. SODA: Spectral Orthogonal Decomposition Adaptation for Diffusion Models
  21. Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Image Generation
  22. Show and Segment: Universal Medical Image Segmentation via In-Context Learning
  23. SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
  24. Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
  25. The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models Via Visual Information Steering
  26. VISIAR: Empower MLLM for Visual Story Ideation
  27. AVID: Any-Length Video Inpainting with Diffusion Model
  28. Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification
  29. BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models
  30. DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
  31. DiMSUM: Diffusion Mamba - A Scalable and Unified Spatial-Frequency Method for Image Generation
  32. Enhanced Deep Unrolled Models Applied to the CMRxRecon2024 Challenge
  33. Generating Enhanced Negatives for Training Language-Based Object Detectors
  34. How to Trace Latent Generative Model Generated Images without Artificial Watermark?
  35. Instantaneous Perception of Moving Objects in 3D
  36. Layout-Agnostic Scene Text Image Synthesis with Diffusion Models
  37. Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate
  38. Learning to Localize Actions in Instructional Videos with LLM-Based Multi-pathway Text-Video Alignment
  39. Rethinking Deep Unrolled Model for Accelerated MRI Reconstruction
  40. SF-V: Single Forward Video Generation Model
  41. Score-Guided Diffusion for 3D Human Recovery
  42. Taming Self-Training for Open-Vocabulary Object Detection