Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA
This story is from 2026-09-22. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.22628v1 Announce Type: new Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the prompt, judge, and scoring pipeline fixed, across four commercial mode…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv cs.AI
Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA