PPaperPicks

Bin Wang

Shanghai Artificial Intelligence Laboratory, Shanghai, China

17 papers at tracked venues · 16 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Joint Knowledge Base Completion and Question Answering by Combining Large Language Models and Small Language Models
  2. MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
  3. MoDora: Tree-Based Semi-Structured Document Analysis System
  4. Chimera: Improving Generalist Model with Domain-Specific Experts
  5. Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
  6. GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
  7. Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
    CVPR 2025 · Bin Wang
  8. LEGION: Learning to Ground and Explain for Synthetic Image Detection
  9. OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
  10. OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
  11. OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
  12. SURVEYFORGE : On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
  13. Distribution-Aware Data Expansion with Diffusion Models
  14. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
  15. OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
  16. Parrot Captions Teach CLIP to Spot Text
  17. VIGC: Visual Instruction Generation and Correction
    AAAI 2024 · Bin Wang