AINewsnow

Updated from 3x3090(2x3090, 1x3090TI) to 2x5090

Upgraded from 3x RTX 3090s to 2x RTX 5090s on my homelab server and picked up a solid speed jump on top of it from a software update (speculative decoding + NVFP4). Setup: llama.cpp (build b11216), running Qwen3.8-27B (abliterated, Q8_0). Blue = old 3090 setup, green = new 5090 setup on the same Q8…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-27 17:04 · r/LocalLLaMA
    Updated from 3x3090(2x3090, 1x3090TI) to 2x5090

More stories

  1. Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding — MarkTechPost
  2. Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine — r/LocalLLM
  3. is switching from llama cpp to vllm worth it — r/LocalLLaMA
  4. Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy — r/LocalLLaMA
  5. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  6. JiRackUltra_1b Runs AI Routing on Any Laptop Without a GPU — AlphaSignal
  7. Advice on models for RAG use case — r/LocalLLaMA
  8. Targeting Applied AI / ML Engineer roles. Need ruthless feedback on my architecture and metrics. — r/deeplearning

Get the daily brief of stories like this at 6:30 every morning →