In Transformer networks why do token embeddings and position embeddings get added?
Hi, going through the Let's Build ChatGPT tutorial here, and prior went through the whole Makemore tutorial that leads up to this tutorial: https://www.youtube.com/watch?v=kCc8FmEb1nY&list=PLAV29EAhk_mX13BqhzdlgM8zkHwpcRajt&index=6&t=2286s When it gets to the point of adding in the attention mechan…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-29 06:17 · r/learnmachinelearning
In Transformer networks why do token embeddings and position embeddings get added?