PPaperPicks

Mohamed Elhoseiny

King Abdullah University of Science and Technology, Thuwal, Saudi Arabia

36 papers at tracked venues · 25 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Sketch2Stitch: GANs for Abstract Sketch-Based Dress Synthesis
  2. Step-by-step Layered Design Generation
  3. XProvence: Zero-Cost Multilingual Context Pruning for Retrieval-Augmented Generation
  4. iMotion-LLM: Instruction-Conditioned Trajectory Generation
  5. 4D-Bench: Benchmarking Multi-Modal Large Language Models for 4D Object Understanding
  6. A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
  7. AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
  8. Aurelia: Test-Time Reasoning Distillation in Audio-Visual LLMs
  9. Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
  10. Diffusion-Based Imaginative Coordination for Bimanual Manipulation
  11. Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents
  12. From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
  13. InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
  14. Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
  15. Local Masked Reconstruction for Efficient Self-Supervised Learning on High-Resolution Images
  16. LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
  17. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
  18. Query-based Knowledge Transfer for Heterogeneous Learning Environments
  19. StoryGPT-V: Large Language Models as Consistent Story Visualizers
  20. Temporal Model-Based Federated Active Medical Image Classification
  21. Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
  22. WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation
  23. AI Art Neural Constellation: Revealing the Collective and Contrastive State of AI-Generated and Human Art
  24. Adversarial Text to Continuous Image Generation
  25. Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
  26. Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation
  27. Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
  28. ImageCaptioner2: Image Captioner for Image Captioning Bias Amplification Assessment
  29. MEERKAT: Audio-Visual Large Language Model for Grounding in Space and Time
  30. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
  31. Multimodal Representation and Retrieval [MRR 2024]
  32. No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
  33. Overcoming Generic Knowledge Loss with Selective Parameter Update
  34. ShapeWalk: Compositional Shape Editing Through Language-Guided Chains
  35. Uni3DL: A Unified Model for 3D Vision-Language Understanding
  36. VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding