AINewsnow

dual 20gb 3080 and 4070ti local llm

Hello just wanted to share my weird setup. I made a huge risk on buying 2 20gb 3080s (modded) a while back and I would like to share what I learned. I am using Hermes agent and llama.cpp and qwen3.8-27b q8 131.1k + vision context on just the 2 3080s. To be frank I'm a complete novice in this field…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-30 02:41 · r/LocalLLM
    dual 20gb 3080 and 4070ti local llm

More stories

  1. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  2. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  3. Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system — r/LocalLLM
  4. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  5. Final-year student in India trying to break into generative-model inference optimization — roadmap feedback? — r/MLQuestions
  6. Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers) — r/LocalLLM
  7. Model Registry (RTX 4090) — r/LocalLLM
  8. Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →