PPaperPicks

Jingqun Tang

15 papers at tracked venues · 14 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
  2. MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
  3. SCORE: Story Coherence and Retrieval Enhancement for AI Narratives
  4. A Bounding Box is Worth One Token - Interleaving Layout and Text in a Large Language Model for Document Understanding
  5. Advancing Sequential Numerical Prediction in Autoregressive Models
  6. Attentive Eraser: Unleashing Diffusion Model's Object Removal Potential via Self-Attention Redirection Guidance
  7. Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
  8. MINDEV: Multi-modal Integrated Diffusion Framework for Video Reconstruction from EEG Signals
  9. MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
    ACL 2025 · Jingqun Tang
  10. OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
  11. ParGo: Bridging Vision-Language with Partial and Global Views
  12. WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
  13. Harmonizing Visual Text Comprehension and Generation
  14. Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer
  15. TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy