Update: Yandex/AliceAI 80B-A3B fine tune progress
loss curve (taken from the last micro of every step, to explain the variation) some help from gemini 3.8 flash high About 40% of the way done with the initial fine tune. The loss is so spiky because I accidentally used the last loss of each micro, rather than the average of each step The training l…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-30 16:30 · r/LocalLLaMA
Update: Yandex/AliceAI 80B-A3B fine tune progress