PPaperPicks

Xiang Bai

52 papers at tracked venues · 41 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at Scale
  2. Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
  3. Doc-V*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
  4. I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing
  5. OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
  6. StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
  7. A Unified Image-Dense Annotation Generation Model for Underwater Scenes
  8. AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
  9. CLIP-AdaM: Adapting Multi-view CLIP for Open-set 3D Object Retrieval
  10. Describe, Adapt and Combine: Empowering CLIP Encoders for Open-Set 3D Object Retrieval
  11. DocThinker: Explainable Multimodal Large Language Models with Rule-Based Reinforcement Learning for Document Understanding
  12. HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
  13. LIRA: Inferring Segmentation in Large Multi-Modal Models with Local Interleaved Region Assistance
  14. LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
  15. MINIMA: Modality Invariant Image Matching
  16. MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
  17. MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
  18. MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
  19. Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
  20. More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
  21. Multi-Scenario Overlapping Text Segmentation with Depth Awareness
  22. NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
  23. OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
  24. Orion: A Holistic End-To-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
  25. PathVG: A New Benchmark and Dataset for Pathology Visual Grounding
  26. PlayerOne: Egocentric World Simulator
  27. Recammaster: Camera-Controlled Generative Rendering From a Single Video
  28. SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
  29. Theorem-Validated Reverse Chain-of-Thought Problem Generation for Geometric Reasoning
  30. Towards Comprehensive Lecture Slides Understanding: Large-Scale Dataset and Effective Method
  31. Training-Free Geometric Image Editing on Diffusion Models
  32. URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model
  33. VIP: Vision Instructed Pre-training for Robotic Manipulation
  34. WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
  35. A Unified Framework for 3D Scene Understanding
  36. Bridging the Gap Between End-to-End and Two-Step Text Spotting
  37. Deciphering Oracle Bone Language with Diffusion Models
  38. Dynamic Adapter Meets Prompt Tuning: Parameter-Efficient Transfer Learning for Point Cloud Analysis
  39. General Object Foundation Model for Images and Videos at Scale
  40. LION: Linear Group RNN for 3D Object Detection in Point Clouds
  41. Make Your ViT-Based Multi-view 3D Detectors Faster via Token Compression
  42. MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks
  43. Monkey: Image Resolution and Text Label are Important Things for Large Multi-Modal Models
  44. OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
  45. OPEN: Object-Wise Position Embedding for Multi-view 3D Object Detection
  46. PSALM: Pixelwise SegmentAtion with Large Multi-modal Model
  47. PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects
  48. PointMamba: A Simple State Space Model for Point Cloud Analysis
  49. SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer
  50. SEED: A Simple and Effective 3D DETR in Point Clouds
  51. WAS: Dataset and Methods for Artistic Text Segmentation
  52. Well Begun is Half Done: The Importance of Initialization in Dataset Distillation