I stopped sending entire PDFs to LLMs.
Here’s why. 👇 A 125-page bank statement can contain ~44,000 tokens of raw text — most of it being letterheads, disclaimers, footers, repeated headers, and data the question never needs. So I built pdfschema , an open-source Python library that extracts only the columns and rows you need . 📉 Token…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-27 13:16 · r/learnmachinelearning
I stopped sending entire PDFs to LLMs.