2x 3090 + 4x P100 — how would you architect this for an agentic RAG system?
I’m building a local AI system for my architecture practice and currently have: 2x RTX 3090 24GB = 48GB 4x Tesla P100 16GB = 64GB 128GB system RAM 2TB NVMe EPYC/PCIe platform So I basically have a 48GB fast pool and a 64GB slower pool . My use case is a large architecture/business database. I have…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-29 18:31 · r/LocalLLM
2x 3090 + 4x P100 — how would you architect this for an agentic RAG system?