cs.CLJul 14, 2026

From Pixels to Pairs: A Comprehensive Benchmark of LLM-Driven Key-Value Extraction in Noisy Document Settings

Authors: Zahra Anvari

Organizations: Independent AI Researcher.

Abstract

Large language models (LLMs) have demonstrated strong capabilities in document key-value pair (KVP) extraction, yet controlled evaluations of their robustness to optical character recognition (OCR) output remain limited. This leaves an important gap in understanding their reliability in real-world OCR-to-LLM pipelines. Unlike end-to-end Vision-Language Models (VLMs), which jointly perform visual perception and semantic extraction, modular pipelines allow these stages and their errors to be isolated and audited. We introduce a controlled benchmark that distinguishes downstream LLM extraction behavior from upstream OCR degradation. It evaluates 136 experimental configurations and 17,688 document-level inferences generated with deterministic decoding across five instruction-tuned open-weight LLMs (2B-8B parameters), three datasets, and four text-quality conditions. The evaluation combines a full zero-shot comparison, targeted one- to three-shot experiments, and a sensitivity analysis of 40 configurations across 20 frozen demonstration sets. By separating Key Recall (annotated-field recovery) from Exact Match and Value F1 (exact and partial value recovery, respectively), we test whether OCR degradation affects field identification and value reproduction differently across models. Our findings challenge three practical assumptions: (1) clean-text performance reliably predicts real-world robustness, (2) model rankings remain consistent across annotation-derived Gold and OCR-derived text, and (3) additional few-shot demonstrations monotonically improve extraction accuracy. The observed model-ranking reversals and unstable few-shot gains expose important reliability risks under noisy document conditions. We release the benchmarking framework, dataset splits, and evaluation scripts to support reproducible research.

Figures & tables

Explore similar work

CardsList
  1. Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction

    Aug 3, 2026Priyashree Roy, Sujitha Martin, Mohammad Rostami +6Confidence EstimationExtraction

  2. Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction

    Sep 14, 2026Kushal Patel, Pushkal Shrivastava, Mackenzie Lees +4Recent Vision-Language ModelsVisual Document Retrieval