I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokens
DISCLAIMER: the post and the model card was made with the assist of AI. Hiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€ Weights, inference code, model card…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-11 10:19 · r/LocalLLaMA
I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokens