From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.17538v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for structured information extraction from documents, yet their behavior under realistic OCR noise remains poorly understood. We present a systematic benchmark of open-source instruction-tuned LLMs fo…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-17 04:00 · arXiv cs.CL
From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings