RTX4090 - Ninfer - Qwen 3.8 27b - 100+ T/S
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
Running the ninfer https://github.com/UDPSendToFailed/ninfer-4090 inference library with an Nvidia 4090 - 24gb with a context of 32k on the neroued\Qwen3.8-27B-NInfer model and I'm getting 100+ t/s and some really good results for an agent driven harness.
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-20 23:04 · r/LocalLLM
RTX4090 - Ninfer - Qwen 3.8 27b - 100+ T/S