I built an open-source PDF extractor for RAG/LLMs that preserves tables, math, and reading order (runs purely on CPU)
Hey guys, While working on NLP research and RAG pipelines, I kept running into the same problem: most PDF parsers completely destroy document structure. Multi-column papers get scrambled, tables become text blobs, and mathematical equations often disappear or become unreadable. I wanted something l…
Read the full story at r/learnmachinelearning ↗
Timeline · 3 reports
- 2026-10-01 03:31 · r/learnmachinelearning
I built an open-source PDF extractor for RAG/LLMs that preserves tables, math, and reading order (runs purely on CPU) - 2026-10-01 03:31 · r/deeplearning
I built an open-source PDF extractor for RAG/LLMs that preserves tables, math, and reading order (runs purely on CPU) - 2026-10-01 03:19 · r/learnmachinelearning
I built an open-source PDF extractor for RAG/LLMs that preserves tables, math, and reading order (runs purely on CPU)