PPaperPicks

Xiaoshuai Sun

35 papers at tracked venues · 32 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval
  2. Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification Approach
  3. ACL: Activating Capability of Linear Attention for Image Restoration
  4. Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
  5. Aigi-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
  6. Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
  7. FlashSloth : Lightning Multimodal Large Language Models via Embedded Visual Compression
  8. HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
  9. IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
  10. MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
  11. Routing Experts: Learning to Route Dynamic Experts in Existing Multi-modal Large Language Models
  12. StoryWeaver: A Unified World Model for Knowledge-Enhanced Story Character Customization
  13. Towards General Visual-Linguistic Face Forgery Detection
  14. γ-MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
  15. 3D-GRES: Generalized 3D Referring Expression Segmentation
  16. 3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression Segmentation
  17. AnyTrans: Translate AnyText in the Image with Large Scale Models
  18. ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
  19. Deep Instruction Tuning for Segment Anything Model
  20. DiffusionFake: Enhancing Generalization in Deepfake Detection via Guided Stable Diffusion
  21. Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
  22. Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
  23. Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
  24. I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing
  25. Improving Panoptic Narrative Grounding by Harnessing Semantic Relationships and Visual Confirmation
  26. Multi-branch Collaborative Learning Network for 3D Visual Grounding
  27. QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual Grounding
  28. RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
  29. Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation
  30. SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
  31. StealthDiffusion: Towards Evading Diffusion Forensic Detection through Diffusion Model
  32. Toward Open-Set Human Object Interaction Detection
  33. Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
  34. X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar Generation
  35. X-RefSeg3D: Enhancing Referring 3D Instance Segmentation via Structured Cross-Modal Graph Neural Networks