Training a tiny 3.6M param byte-level Transformer for code generation from scratch on CPU. What should be my next steps?
Hi everyone, I'm working on a personal learning project called MOTANAXY. The goal is to build and train a tiny causal Transformer from scratch (random weights) specifically for Python code generation. I'm currently training entirely on CPU. Current Architecture (v2): Type: Byte-level causal languag…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-10-02 09:10 · r/learnmachinelearning
Training a tiny 3.6M param byte-level Transformer for code generation from scratch on CPU. What should be my next steps?