AINewsnow

I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs

No transformers, no Trainer. Hand-written BPE tokenizer, Llama-style decoder, DDP training loop, SFT, all in plain PyTorch. 20B tokens on 2x RTX A4500, six days, about $21 of electricity. Base model gets 13.4% on HumanEval, the chat version 26.8% (Codex-12B was 28.8%). It is bad at everything that…

Read the full story at r/learnmachinelearning ↗

Timeline · 2 reports

  1. 2026-09-21 23:12 · r/deeplearning
    I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs
  2. 2026-09-21 23:11 · r/learnmachinelearning
    I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs

More stories

  1. Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
  2. Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF) — r/LocalLLM
  3. Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken — r/LocalLLaMA
  4. PXA v2026.09.20 — my inference engine for old Teslas (P100 / V100 / 1080 Ti): Gemma 4 MoE, tensor split on by default, and ahead of stock llama.cpp on every cell on my rig — r/LocalLLM
  5. Who's getting above 50 tok/s on AMD 9070, R9700 GPUs? — r/LocalLLM
  6. Ternary Bonsai 2 (27B) fails to load in LM Studio and oMLX. I made fixes for both (GGUF PQ2_0/PTQ1_0 + MLX 2-bit) — r/LocalLLM
  7. You can use any LLM just like JEV — r/LocalLLaMA
  8. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →