Swift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090
Hi everyone :) The amazing Swift finetunes of Qwen3.8 27B generate much fewer tokens at mostly similar benchmark performance to the original model, while repos like HyperQwen (formerly syv-ai/qwen38-27b-rtx3090) deliver insane TPS on an RTX 3090. To get the best of both worlds, I adapted Swift 1 an…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-28 20:52 · r/LocalLLaMA
Swift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090