PSSA, a plastic state space model, beats a parameter-matched transformer on held-out text and generates ~12x faster on CPU
I built a from-scratch architecture called PSSA (plastic state space architecture) and trained it against a parameter-matched transformer baseline on the same corpus, same 12.7M tokens, same tokenizer and schedule. Held-out results on a 198,939-token slice neither run saw: cross-entropy 3.997 vs 4.…
Read the full story at r/deeplearning ↗
Timeline · 3 reports
- 2026-09-30 00:31 · r/learnmachinelearning
PSSA, a plastic state space model, beats a parameter-matched transformer on held-out text and generates ~12x faster on CPU - 2026-09-29 23:31 · r/artificial
PSSA, a plastic state space model, beats a parameter-matched transformer on held-out text and generates ~12x faster on CPU - 2026-09-29 23:07 · r/deeplearning
PSSA, a plastic state space model, beats a parameter-matched transformer on held-out text and generates ~12x faster on CPU