PPaperPicks

Yuliang Liu

23 papers at tracked venues · 20 at CORE A* · active 20242025

Venues

Frequent coauthors

Papers

  1. AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
    ICML 2025 · Yuliang Liu
  2. DocThinker: Explainable Multimodal Large Language Models with Rule-Based Reinforcement Learning for Document Understanding
  3. LIRA: Inferring Segmentation in Large Multi-Modal Models with Local Interleaved Region Assistance
  4. LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
  5. MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
  6. MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
  7. Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
  8. Multi-Scenario Overlapping Text Segmentation with Depth Awareness
  9. OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
  10. SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
  11. Theorem-Validated Reverse Chain-of-Thought Problem Generation for Geometric Reasoning
  12. Towards Comprehensive Lecture Slides Understanding: Large-Scale Dataset and Effective Method
  13. Training-Free Geometric Image Editing on Diffusion Models
  14. WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
  15. AP-Adapter: Improving Generalization of Automatic Prompts on Unseen Text-to-Image Diffusion Models
  16. Bridging the Gap Between End-to-End and Two-Step Text Spotting
  17. Deciphering Oracle Bone Language with Diffusion Models
  18. MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks
  19. Monkey: Image Resolution and Text Label are Important Things for Large Multi-Modal Models
  20. OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
  21. ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
  22. Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
  23. Well Begun is Half Done: The Importance of Initialization in Dataset Distillation