Is there still strong interest in a dense 9b model?
This story is from 2026-09-13. It is preserved in the archive; the latest stories are on the live feed.
I have a full model, it's ready to train. It's ~9b parameters. 9.4b to be more exact. That includes a 1/2/3 Engram table, Moonshot's AttnRes modeling, and RoPE / NoPE layering at 3:1 as more or less validated by most major labs. It uses the Llama 3 series tokenizer and LM Head as an initial start.…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-13 06:08 · r/LocalLLaMA
Is there still strong interest in a dense 9b model?