AINewsnow

Update: Yandex/AliceAI 80B-A3B fine tune progress

loss curve (taken from the last micro of every step, to explain the variation) some help from gemini 3.8 flash high About 40% of the way done with the initial fine tune. The loss is so spiky because I accidentally used the last loss of each micro, rather than the average of each step The training l…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-30 16:30 · r/LocalLLaMA
    Update: Yandex/AliceAI 80B-A3B fine tune progress

More stories

  1. Gemini 4 Argon: our next era of frontier intelligence — Google DeepMind Blog
  2. Google rolls out Gemini 4 Argon to a small group of cybersecurity partners and says it outperforms GPT-6 Astra on certain coding and knowledge work benchmarks (Madison Mills/Axios) — Techmeme
  3. Gemini 4 Argon — Hacker News Front Page
  4. Google announces Gemini 4 Argon AI model, but you can't use it yet — Ars Technica AI
  5. Google Releases New Gemini Model With Guardrails Amid A.I. Safety Debate — New York Times Technology
  6. Let skills in Gemini tackle your most repetitive tasks — Google Gemini Blog
  7. See what 4 builders are making with Gemini 3.8 Flash — Google Gemini Blog
  8. 3 ways this grocer cooks for 200 guests with Gemini — Google Gemini Blog

Get the daily brief of stories like this at 6:30 every morning →