PPaperPicks

Fahad Shahbaz Khan

Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, AUE

81 papers at tracked venues · 56 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. A Multi-Agent Diffusion Approach for MRI Anomaly Segmentation via Modality-Specific LoRA Specialization
  2. AURORA: Augmented Understanding via Structured Reasoning and Reinforcement Learning for Reference Audio-Visual Segmentation
  3. Bring Your Dreams to Life: Continual Text-to-Video Customization
  4. GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support
  5. MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
  6. Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework
  7. A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
  8. ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions
  9. AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and Modulation
  10. AirCast: Improving Air Pollution Forecasting Through Multi-Variable Data Alignment
  11. All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages
  12. Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
  13. BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalities
  14. CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
  15. DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
  16. DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
  17. EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues
  18. Efficient Video Object Segmentation via Modulated Cross-Attention Memory
  19. Enhancing Novel Object Detection via Cooperative Foundational Models
  20. GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
  21. GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
  22. GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
  23. GroupMamba: Efficient Group-Based Visual State Space Model
  24. Hierarchical Self-supervised Adversarial Training for Robust Vision Models in Histopathology
  25. Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
  26. How Good is my Video-LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
  27. InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration
  28. Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
  29. KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
  30. LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
  31. LawDIS: Language-Window-Based Controllable Dichotomous Image Segmentation
  32. LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
  33. MAviS: A Multimodal Conversational Assistant For Avian Species
  34. MPromer: A Unified Diffusion-Based Framework for Scalable and Generalizable Multi-Modal Medical Image Segmentation
  35. One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
  36. One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
  37. Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation
    ICLR 2025 ·
    Mohamed El Amine Boudjoghra
  38. Open3DTrack: Towards Open-Vocabulary 3D Multi-Object Tracking
  39. Palo: A Polyglot Large Multimodal Model for 5B People
  40. RAGNet: Large-Scale Reasoning-Based Affordance Segmentation Benchmark Towards General Grasping
  41. SPARTA: Spectral Prompt Agnostic Adversarial Attack on Medical Vision-Language Models
  42. TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
  43. Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts
  44. Towards Evaluating the Robustness of Visual State Space Models
  45. VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
  46. VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
  47. Vocabulary-Free Fine-Grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
  48. ZeroDiff: Solidified Visual-semantic Correlation in Zero-Shot Learning
  49. BAPLe: Backdoor Attacks on Medical Foundational Models Using Prompt Learning
  50. BiMediX: Bilingual Medical Mixture of Experts LLM
  51. Bidirectional Reciprocative Information Communication for Few-Shot Semantic Segmentation
  52. CONDA: Condensed Deep Association Learning for Co-salient Object Detection
  53. Composed Video Retrieval via Enriched Context and Discriminative Embeddings
  54. Continual Learning and Unknown Object Discovery in 3D Scenes via Self-distillation
    ECCV 2024 ·
    Mohamed El Amine Boudjoghra
  55. Cross-Modal Self-Training: Aligning Images and Pointclouds to learn Classification without Labels
  56. DB-SAM: Delving into High Quality Universal Medical Image Segmentation
  57. Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
  58. GLaMM: Pixel Grounding Large Multimodal Model
  59. GeoChat: Grounded Large Vision-Language Model for Remote Sensing
  60. Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
  61. Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
  62. How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?
  63. Learning Camouflaged Object Detection from Noisy Pseudo Label
  64. Long-Tailed 3D Semantic Segmentation with Adaptive Weight Constraint and Sampling
  65. MaskFactory: Towards High-quality Synthetic Data Generation for Dichotomous Image Segmentation
  66. MedContext: Learning Contextual Cues for Efficient Volumetric Medical Segmentation
  67. Modulate Your Spectrum in Self-Supervised Learning
  68. On Evaluating Adversarial Robustness of Volumetric Medical Segmentation Models
  69. Open-Vocabulary Temporal Action Localization using Multimodal Guidance
  70. Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
  71. Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
  72. S3A: Towards Realistic Zero-Shot Classification via Self Structural Semantic Alignment
  73. SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation
  74. Self-Distilled Masked Auto-Encoders are Efficient Video Anomaly Detectors
  75. Semi-supervised Open-World Object Detection
  76. Sentence-level Prompts Benefit Composed Image Retrieval
  77. Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
  78. VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt Learning
  79. Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
  80. VideoGrounding-DINO: Towards Open-Vocabulary Spatio- Temporal Video Grounding
  81. Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot Learning