AINewsnow

Transformer becomes catastrophically ill-conditioned after a few tiny parameter updates: same batch goes from grad norm 5.6 → 89 while weights move <0.1%/step. What mechanism could cause this?

My Transformer trains normally for some steps, then begins to enter a parameter state where backpropagation through the middle/lower layers magnifies gradients massively, despite the forward activations and weights being normal. Eventually, the gradients explode, get clipped globally, and learning…

Read the full story at r/deeplearning ↗

Timeline · 1 report

  1. 2026-09-26 21:10 · r/deeplearning
    Transformer becomes catastrophically ill-conditioned after a few tiny parameter updates: same batch goes from grad norm 5.6 → 89 while weights move <0.1%/step. What mechanism could cause this?

More stories

  1. Introducing Gemini 3.8 Live with Live Avatar — Google Gemini Blog
  2. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  3. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times AI
  4. An OpenAI agent hacked Medicare. Will anyone be held responsible? — The Conversation AI (US)
  5. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  6. Am I the only one who actually likes GPT-6 Sol and Luna? — r/ChatGPT
  7. Is Qwen Flash Next at like Q2 better than 27B at Q4? — r/LocalLLaMA
  8. Appeals Court Lets the Pentagon Designate Anthropic a Supply-Chain Risk — Wired AI

Get the daily brief of stories like this at 6:30 every morning →