AINewsnow

A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass

arXiv:2610.03940v1 Announce Type: new Abstract: During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and th…

Read the full story at arXiv cs.CL ↗

Timeline · 1 report

  1. 2026-10-06 04:00 · arXiv cs.CL
    A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass

More stories

  1. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  2. Introducing GLM 5.3 on Amazon Bedrock — AWS Machine Learning Blog
  3. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  4. OpenAI safety employee resigns, claiming the company’s ‘culture is broken’ — TechCrunch AI
  5. Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 — r/LocalLLaMA
  6. Supercharge regulated workloads with Claude Code and Amazon Bedrock — AWS Machine Learning Blog
  7. can i run qwen flash next with these specs, or am i out of luck? — r/LocalLLM
  8. The Story of Qwen: Alibaba's AI Models From 7B to 2.4T — MarkTechPost

Get the daily brief of stories like this at 6:30 every morning →