Nova v3: 148M model, 1% of the data, same benchmark scores as SmolLM2-135M
I trained Nova v3, a 148M parameter chat model, from scratch. It matches SmolLM2-135M on HellaSwag, ARC-Easy, PIQA, and WinoGrande while being trained on about 1 percent of the tokens SmolLM2 saw. Model: huggingface.co/plasmova/nova-v3 Architecture and tokenizer details, training data composition,…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-29 19:36 · r/learnmachinelearning
Nova v3: 148M model, 1% of the data, same benchmark scores as SmolLM2-135M