PPaperPicks

Xiangtai Li

51 papers at tracked venues · 45 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud Classification
  2. AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
  3. Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
  4. Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language
  5. Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
  6. Bridge Feature Matching and Cross-Modal Alignment with Mutual-Filtering for Zero-Shot Anomaly Detection
  7. Conditional Panoramic Image Generation via Masked Autoregressive Modeling
  8. Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
  9. DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
  10. DreamRelation: Bridging Customization and Relation Generation
  11. Explore In-Context Segmentation via Latent Diffusion Models
  12. Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene
  13. MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
  14. Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
  15. MelodyEdit: Zero-shot Music Editing with Disentangled Inversion Control
  16. NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results
  17. OmniAudio: Generating Spatial Audio from 360-Degree Video
  18. On Path to Multimodal Generalist: General-Level and General-Bench
  19. Point Cloud Mamba: Point Cloud Learning via State Space Model
  20. PointDGMamba: Domain Generalization of Point Cloud Classification via Generalized State Space Model
  21. PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning
  22. QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video Editing
  23. RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
  24. RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection
  25. SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model
  26. The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
  27. Three-Dimensional Trajectory Prediction with 3DMoTraj Dataset
  28. Towards Semantic Equivalence of Tokenization in Multimodal LLM
  29. UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
  30. Unified Dense Prediction of Video Diffusion
  31. VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
  32. BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model
  33. CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
  34. DGMamba: Domain Generalization via Generalized State Space Model
  35. Face-Adapter for Pre-trained Diffusion Models with Fine-Grained ID and Attribute Control
  36. From Multimodal LLM to Human-level AI: Modality, Instruction, Reasoning and Beyond
  37. GenView: Enhancing View Quality with Pretrained Generative Model for Self-Supervised Learning
  38. Improving Video Segmentation via Dynamic Anchor Queries
  39. MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly Detection
  40. MotionBooth: Motion-Aware Customized Text-to-Video Generation
  41. OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
  42. OMG-Seg: Is One Model Good Enough for all Segmentation?
    CVPR 2024 · Xiangtai Li
  43. Open-Vocabulary SAM: Segment and Recognize Twenty-Thousand Classes Interactively
  44. RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation
  45. Referring Image Editing: Object-Level Image Editing via Referring Expressions
  46. SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified Flow
  47. Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning
  48. Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-image
  49. Towards Language-Driven Video Inpainting via Multimodal Large Language Models
  50. VG4D: Vision-Language Model Goes 4D Video Recognition
  51. You Can't Ignore Either: Unifying Structure and Feature Denoising for Robust Graph Learning