AINewsnow

2x 3090 + 4x P100 — how would you architect this for an agentic RAG system?

I’m building a local AI system for my architecture practice and currently have: 2x RTX 3090 24GB = 48GB 4x Tesla P100 16GB = 64GB 128GB system RAM 2TB NVMe EPYC/PCIe platform So I basically have a 48GB fast pool and a 64GB slower pool . My use case is a large architecture/business database. I have…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-29 18:31 · r/LocalLLM
    2x 3090 + 4x P100 — how would you architect this for an agentic RAG system?

More stories

  1. 2x Tesla p100s, q6_k quant, Qwen 3.8 27B ~60tps V3.0 — r/LocalLLaMA
  2. Don't trust frontier models when asking about budget hardware! — r/LocalLLaMA
  3. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  4. The Future Is for Everyone: Muse for Small Business — Meta Newsroom
  5. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  6. OpenAI launches Dots, its Muse competitor — The Verge AI
  7. OpenAI pauses AI training, launches ‘extensive’ review after multiple rogue agent incidents — Mint AI
  8. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →