I trained a 125M GPT on a free Colab T4: 655k tokens in 83 seconds, and the loss looks broken until you check the init
I wanted a real number for "can you actually train a small LM on the free tier", so I ran one end to end this morning and wrote down what the runtime printed instead of guessing. Free Colab T4 (Tesla T4, 14.56 GB, 2 vCPU). A 124,439,808-parameter GPT-2-shaped model, context 1024, micro-batch 1 with…
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-09-24 12:35 · r/deeplearning
I trained a 125M GPT on a free Colab T4: 655k tokens in 83 seconds, and the loss looks broken until you check the init
More stories
- Introducing GPT-6 Sol and Luna — OpenAI News
- OpenAI Agent Hacked Australian Government Website — Wall Street Journal Technology
- Meet the Data Agent in ChatGPT Work — OpenAI YouTube
- Anthropic and OpenAI roll out cheaper models in first release since call for slowdown — CNBC Technology
- Anthropic launches Claude Opus 5.5, promising Fable-level performance at a lower price — Mashable AI
- Deploy and manage coding agents at scale with the Unity Gateway CLI — Databricks Blog
- New and need help, — r/comfyui
- Harvey turns legal context into stronger drafts with GPT-6 Astra — OpenAI News
Get the daily brief of stories like this at 6:30 every morning →