LoRA over GGUF: Train Qwen3.8-Flash-Next in 40G VRAM
https://github.com/woct0rdho/transformers5-qwen3.5-recipe An update to my LoRA over GGUF series: Now we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM, with no CPU offloading, with engram on disk that does not reduce training speed. On Strix Halo it trains context chunk size 20…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-09 03:54 · r/LocalLLaMA
LoRA over GGUF: Train Qwen3.8-Flash-Next in 40G VRAM