AINewsnow

I trained a 1.46M parameter language model on CPU — then discovered 26.82% validation leakage

This story is from 2026-08-28. It is preserved in the archive; the latest stories are on the live feed.

I've been building COLLISION-1.46M, a decoder-only Transformer trained completely from scratch on my laptop CPU. First serious run: • 1,462,464 parameters • 2.4M training tokens • CPU-only • Custom BPE tokenizer Phase 5 validation perplexity: 62.86 Instead of immediately making the model bigger, I…

Read the full story at r/learnmachinelearning ↗

Timeline · 1 report

  1. 2026-08-28 13:53 · r/learnmachinelearning
    I trained a 1.46M parameter language model on CPU — then discovered 26.82% validation leakage

More stories

  1. AI agents/automation suggestions for a solo biz — r/AI_Agents
  2. Getting more accurate results - personalizations — r/ArtificialInteligence
  3. Opti 27B: Qwen3.8-27B in 11.8 GB at 3.47 bpw, within 0.5% of FP16 perplexity and matching Q4_K_M at 30% fewer bytes. Patched llama.cpp runtime, source public, reproduce with one command — r/LocalLLM
  4. Google's Gemini AI hacks three other companies during security test — Sky News Technology
  5. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  6. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  7. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  8. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology

Get the daily brief of stories like this at 6:30 every morning →