Introducing Infernix - much faster than Strata on a 5090!
I've spent the last week working on my inference engine and it now runs Qwen3.8-Flash-Next at a real usable quant (NVIDIA's NVFP4 model) significantly faster than Strata running Unsloth's UD-Q4_K_XL quant. On my agentic workflow benchmark: Metric Strata Infernix Change Average time to first token 3…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-09 04:34 · r/LocalLLM
Introducing Infernix - much faster than Strata on a 5090!