I trained my own 150M non-Transformer language model from scratch on 300M tokens — WarpState
This story is from 2026-08-28. It is preserved in the archive; the latest stories are on the live feed.
Hi everyone, I’ve been experimenting with alternative language-model architectures for a while, and I recently finished the first complete pretraining run of a new architecture I’m calling WarpState . This is still an experimental proof of concept, not a claim that it beats Transformers or existing…
Read the full story at r/OpenAI ↗
Timeline · 1 report
- 2026-08-28 21:41 · r/OpenAI
I trained my own 150M non-Transformer language model from scratch on 300M tokens — WarpState