In Transformer why are attention block weights and feed forward block weights optimized in same optimization run?
Hi, still going through the Let's Make ChatGPT tutorial here: https://youtu.be/kCc8FmEb1nY?list=PLAV29EAhk_mX13BqhzdlgM8zkHwpcRajt&t=5158 In the video at time shown we hear about how the attention layer captures one level of meaning, and how once that meaning is captured, it is sent through feed fo…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-29 17:31 · r/learnmachinelearning
In Transformer why are attention block weights and feed forward block weights optimized in same optimization run?