AINewsnow

I built Ninfer 4080 for 16GB class GPUs

Hi everyone, TL/DR I created NInfer 4080 to run ISTA-DASLab-Qwen-3.8-27B-GSQ at 100k context on an RTX 4080 16GB GPU using way more of the hardware capabilities ( max overall: 2720 tok/s prefill, 262 tok/s generation ) and sharing it with the community now so others can also have the benefit. https…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-10-03 18:59 · r/LocalLLM
    I built Ninfer 4080 for 16GB class GPUs
  2. 2026-10-03 18:58 · r/LocalLLaMA
    I built Ninfer 4080 for 16GB class GPUs

More stories

  1. Qwen Flash Next MTP work restarted — r/LocalLLaMA
  2. Viggle turbo v0.3 for Qwen image 2.1: less grain and cleaner surfaces — r/StableDiffusion
  3. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  4. Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp — r/LocalLLaMA
  5. Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it — r/machinelearningnews
  6. I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. — r/LocalLLaMA
  7. What is your experience with bonsai 2 27b? — r/ArtificialInteligence
  8. Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print. — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →