PPaperPicks

Lianwen Jin

42 papers at tracked venues · 33 at CORE A* · active 20242026

Venues

Frequent coauthors

Papers

  1. Document Analysis and Recognition - ICDAR 2025 Workshops - Wuhan, China, September 20-21, 2025, Proceedings, Part I
    ICDAR 2026 · Lianwen Jin
  2. Document Analysis and Recognition - ICDAR 2025 Workshops - Wuhan, China, September 20-21, 2025, Proceedings, Part II
    ICDAR 2026 · Lianwen Jin
  3. Draft, Verify, Restore: Self-Refining Historical Inscription Restoration with a Unified MLLM
  4. Frequency Mining Empowered by Text Aggregation: A New Perspective on Document Image Tampering Detection
  5. HisDoc-OCR: Restoring Visual Grounding in MLLMs for Chinese Historical Document OCR
  6. PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography
  7. TextShield-R1: Reinforced Reasoning for Tampered Text Detection
  8. URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding
  9. CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
  10. Capturing More: Learning Multi-Domain Representations for Robust Online Handwriting Verification
  11. DevInSight: Weaving Path Development Into Online Signature Verification
  12. DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
  13. DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
  14. From Pixels to Semantics: A Novel MLLM-Driven Approach for Explainable Tampered Text Detection
  15. Hallucination-Aware Prompt Optimization for Text-to-Video Synthesis
  16. HisDoc-DETR: Integrating Semantic Learning and Feature Fusion for Historical Document Layout Analysis
  17. Large-Scale Corpus Construction and Retrieval-Augmented Generation for Ancient Chinese Poetry: New Method and Data Insights
  18. MCCD: A Multi-attribute Chinese Calligraphy Character Dataset Annotated with Script Styles, Dynasties, and Calligraphers
  19. MCS-Bench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in Chinese Classical Studies
  20. Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
  21. OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
  22. Predicting the Original Appearance of Damaged Historical Documents
  23. RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
  24. Revisiting Tampered Scene Text Detection in the Era of Generative AI
  25. Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration
  26. TongGu-VL: Advancing Visual-Language Understanding in Chinese Classical Studies through Parameter Sensitivity-Guided Instruction Tuning
  27. Bridging the Gap Between End-to-End and Two-Step Text Spotting
  28. Deciphering Oracle Bone Language with Diffusion Models
  29. DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
  30. DocNLC: A Document Image Enhancement Framework with Normalized and Latent Contrastive Representation for Multiple Degradations
  31. DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
  32. FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive Learning
  33. M2Doc: A Multi-Modal Fusion Approach for Document Layout Analysis
  34. PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction
  35. PPTSER: A Plug-and-Play Tag-guided Method for Few-shot Semantic Entity Recognition on Visually-rich Documents
  36. RDLNet: A Novel and Accurate Real-world Document Localization Method
  37. TongGu: Mastering Classical Chinese Understanding with Knowledge-Grounded Large Language Models
  38. Towards Modern Image Manipulation Localization: A Large-Scale Dataset and Novel Methods
  39. UPOCR: Towards Unified Pixel-Level OCR Interface
  40. ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
  41. VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
  42. WenMind: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Classical Literature and Language Arts