Qwen 3.8 + n-gram
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
Has anyone considered some architectural edits to the 27b model, such as: Implementation of LoRA/QLoRA on selected layers; training a modest sized n-gram table and adding the integration layers? Seems like a good project. ESP given how it seems to be the current star of local hosted models! It’s th…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-07 18:43 · r/LocalLLM
Qwen 3.8 + n-gram