Document Image Understanding

Momentum

7 papers in the last four weeks, level with the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 92

All topics
CardsList
  1. OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents

    Jul 10, 2026Yang Chen, Yunwen Li, Yufan Shen +6VLM EvaluationDocument Understanding

  2. LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding

    Jul 7, 2026Guang Yang, Brian Siyuan Zheng, Victoria Ebert +1Optical Music RecognitionAutomatic Music Transcription

  3. HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

    Jul 6, 2026Gengluo Li, Xingyu Wan, Shangpin Peng +20Vision-Language ModelsEfficient VLM Inference

  4. ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation

    Jul 4, 2026Enshuo Hsu, Jin Zhou, Kirk RobertsVLM EvaluationOptical Character Recognition

  5. A Vision Based System for Guided and Collaborative Reconstruction of Fragmented Documents

    Jul 3, 2026Oliver Krumpek, Diana LeoHuman-Robot CollaborationCultural Heritage

  6. SINA: A Fully Automated Circuit Schematic Image to Netlist Generator Using Artificial Intelligence

    Jul 2, 2026Saoud Aldowaish, Yashwanth Karumanchi, Kai-Chen Chiang +4Electronic Design AutomationDocument Image Understanding

  7. DEMUN: Fast and accurate discovery of music notation in very large collections

    Jun 30, 2026Vojtěch Dvořák, Filip Bím, Jiří Mayer +5Optical Music RecognitionVisual Document Retrieval

  8. Multimodal Graph RAG for Long-range Visually Rich Document Understanding

    Jun 27, 2026Yi-Cheng Wang, Chu-Song ChenMultimodal RAGVisual Question Answering

  9. Joint Transcription and Decryption of Images of Encrypted Handwritten Documents: A Comparison with the Traditional Pipeline

    Jun 26, 2026Marino Oliveros-Blanco, Lei Kang, Alicia Fornés +1Optical Character RecognitionDocument Image Understanding

  10. An LMM for Precisely Grounding Elements in Documents

    Jun 23, 2026Yijian Lu, Chuangxin Zhao, Kai Sun +3Visual ReasoningLarge Vision-Language Models

  11. DR-Mamba: Automatic Inference-Time Domain Adaptation for Document Image Binarization via Sample-Conditioned Detail-Background Suppression

    Jun 21, 2026Sheng-Wei Chan, Jen-Shiun ChiangMambaTest-Time Adaptation

  12. Text region detection in historical astronomical diagrams

    Jun 14, 2026Zeynep Sonat Baltacı, Raphaël Baena, Fei Meng +4Document UnderstandingHistorical Document Analysis

  13. IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products

    Jun 12, 2026Haonan Qi, Jin Cao, Yongqi Zhang +9Multimodal Large Language ModelsMultimodal Model Evaluation

  14. ParseFixer: An Agentic Framework for Document Parsing via Selective Multimodal Correction

    Jun 10, 2026LeKai Yu, Hao Liu, Kun Wang +4Document ParsingDocument Layout Analysis

  15. ERN-Net : Evolving Reason Node-Net for Document Binarization

    Jun 10, 2026Hsin-Jui Pan, Sheng-Wei Chan, Jen-Shiung ChiangConvolutional Neural NetworksDocument Image Understanding

  16. MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

    Jun 8, 2026David Setiawan, Temuulen Khishigsuren, Milind Agarwal +3Low-Resource Language ProcessingOptical Character Recognition

  17. Intelligent Character Recognition of Handwritten Forms with Deep Neural Networks

    Jun 7, 2026Hartwig GrabowskiOptical Character RecognitionDocument Image Understanding

  18. DeepMine-Mamba: Mitigating Information Dilution in Mamba-Based State Space Models for Document Image Binarization

    Jun 7, 2026Sheng-Wei Chan, Yung-Che Wang, Hsin-Jui Pan +2MambaDocument Image Understanding

  19. End-to-End Text Line Detection and Ordering

    Jun 2, 2026Benjamin KiesslingDocument Layout AnalysisDocument Image Understanding

  20. Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

    May 31, 2026Minglai Yang, Xinyan Velocity Yu, Pengyuan Li +22VLM EvaluationDocument Parsing

  21. HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

    May 31, 2026Issa Sugiura, Shuhei Kurita, Yusuke Oda +1VLM EvaluationData Visualization

  22. ABot-OCR Technical Report

    May 27, 2026Kaitao Jiang, Ruiyan Gong, Xiaolong Cheng +3Vision-Language ModelsDocument Parsing

  23. Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing

    May 25, 2026Kateryna Lutsai, Dana Křivánková, Pavel Straňák +1Document UnderstandingHistorical Document Analysis

  24. FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers

    May 21, 2026Laziz Hamdi, Amine Tamasna, Pascal Boisson +1Table Structure RecognitionDocument Image Understanding

  25. MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing

    May 21, 2026Bangbang Zhou, Hangdi Xing, Yifan Chen +8Document ParsingDocument Layout Analysis