PPaperPicks

Jiayi Ji

36 papers at tracked venues · 33 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. 3D-DRES: Detailed 3D Referring Expression Segmentation
  2. CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval
  3. FIND: A Simple Yet Effective Baseline for Diffusion-Generated Image Detection
  4. QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension
  5. Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification Approach
  6. ACL: Activating Capability of Linear Attention for Image Restoration
  7. Aigi-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
  8. DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehension
  9. GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification
  10. HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
  11. IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
  12. Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive Segmentation
  13. JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
  14. MDReID: Modality-Decoupled Learning for Any-to-Any Multi-Modal Object Re-Identification
  15. MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
  16. Multi-Modal Object Re-identification via Sparse Mixture-of-Experts
  17. Towards General Visual-Linguistic Face Forgery Detection
  18. Towards Semantic Equivalence of Tokenization in Multimodal LLM
  19. Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
  20. γ-MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
  21. 3D-GRES: Generalized 3D Referring Expression Segmentation
  22. 3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression Segmentation
  23. APL: Anchor-Based Prompt Learning for One-Stage Weakly Supervised Referring Expression Comprehension
  24. ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
  25. Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
  26. Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
  27. I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing
  28. Improving Panoptic Narrative Grounding by Harnessing Semantic Relationships and Visual Confirmation
  29. Multi-branch Collaborative Learning Network for 3D Visual Grounding
  30. RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
  31. Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation
  32. SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
  33. Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-image
  34. Toward Open-Set Human Object Interaction Detection
  35. X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar Generation
  36. X-RefSeg3D: Enhancing Referring 3D Instance Segmentation via Structured Cross-Modal Graph Neural Networks