scaled sigmoid to bound log-decay in kimi K3 KDA
https://preview.redd.it/jzxao82t8fuh1.png?width=1996&format=png&auto=webp&s=7391b767ab01c2faa9ef09753824caecff889f56 i've been trying to gain some thorough intuition about `g_min` bound, which the kimi team used for K3 in kda. i i think i get that the higher order intuition here is to not let the r…
Read the full story at r/learnmachinelearning ↗
Timeline · 2 reports
- 2026-10-09 11:08 · r/deeplearning
scaled sigmoid to bound log-decay in kimi K3 KDA - 2026-10-09 11:07 · r/learnmachinelearning
scaled sigmoid to bound log-decay in kimi K3 KDA