AINewsnow

In Transformer why are attention block weights and feed forward block weights optimized in same optimization run?

Hi, still going through the Let's Make ChatGPT tutorial here: https://youtu.be/kCc8FmEb1nY?list=PLAV29EAhk_mX13BqhzdlgM8zkHwpcRajt&t=5158 In the video at time shown we hear about how the attention layer captures one level of meaning, and how once that meaning is captured, it is sent through feed fo…

Read the full story at r/learnmachinelearning ↗

Timeline · 1 report

  1. 2026-09-29 17:31 · r/learnmachinelearning
    In Transformer why are attention block weights and feed forward block weights optimized in same optimization run?

More stories

  1. OpenAI launches Dots, its Muse competitor — The Verge AI
  2. OpenAI pauses AI training, launches ‘extensive’ review after multiple rogue agent incidents — Mint AI
  3. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
  4. OpenAI Scraps Release of New AI Model Over Safety Concerns — Wall Street Journal Technology
  5. OpenAI pulls the plug on GPT 6.1 Astra as agents keep crossing lines — InfoWorld AI
  6. Introducing GPT-6.1 Sol — OpenAI News
  7. OpenAI Will Keep Driving Down AI Prices, Sam Altman Says — Bloomberg AI
  8. GPT-6 SOL AND LUNA ARE OUT!!! — Matthew Berman

Get the daily brief of stories like this at 6:30 every morning →