One 128 GB Strix Halo box running a full offline RAG stack (8B embedder + 8B reranker + 122B-A10B answerer + 70B judge): what worked and what I'd change
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
Sharing the setup because most "can a 128 GB APU do real work" threads end in speculation. This is a fully offline RAG over 27 technical books (~11k pages, 32k chunks), one machine, AMD Strix Halo, 128 GB unified memory, Ollama + embedded Qdrant. No cloud anywhere in the pipeline. Who does what Emb…
Read the full story at r/LocalLLM ↗