Do Transformer representations progressively structure across depth and time? Results from 8 open models
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
Hi everyone, I’ve just published a new preprint that brings together several months of experiments on hidden-state dynamics in small open Transformer models. The question is fairly simple: During inference, do internal representations simply change from layer to layer, or is there evidence of a mor…
Read the full story at r/deeplearning ↗
Timeline · 3 reports
- 2026-08-27 10:37 · r/PromptEngineering
Do Transformer representations progressively structure across depth and time? Results from 8 open models - 2026-08-26 18:56 · r/learnmachinelearning
Do Transformer representations progressively structure across depth and time? Results from 8 open models - 2026-08-26 18:54 · r/deeplearning
Do Transformer representations progressively structure across depth and time? Results from 8 open models