AINewsnow

scaled sigmoid to bound log-decay in kimi K3 KDA

https://preview.redd.it/jzxao82t8fuh1.png?width=1996&format=png&auto=webp&s=7391b767ab01c2faa9ef09753824caecff889f56 i've been trying to gain some thorough intuition about `g_min` bound, which the kimi team used for K3 in kda. i i think i get that the higher order intuition here is to not let the r…

Read the full story at r/learnmachinelearning ↗

Timeline · 2 reports

  1. 2026-10-09 11:08 · r/deeplearning
    scaled sigmoid to bound log-decay in kimi K3 KDA
  2. 2026-10-09 11:07 · r/learnmachinelearning
    scaled sigmoid to bound log-decay in kimi K3 KDA

More stories

  1. China’s open-weight AI models are winning global users. Who is capturing the value? — South China Morning Post Tech
  2. Quantization of Linear-Attention (Qwen & Kimi) — r/LocalLLaMA
  3. kimi K3 Mental BreakDown — r/ArtificialInteligence
  4. How to fine tune a model ? — r/LocalLLM
  5. Goodfire Cuts AI Jailbreaks From 66 to Zero on Kimi K3 — AlphaSignal
  6. Goodfire Deploys Probe-Based Cyber Monitors for Kimi K3 and GLM 5.3 — Unite.AI
  7. Could 4 M5 Ultra, 512 gb Mac Studios Run Kimi K3? How Many TPS Would it Get? — r/LocalLLM
  8. Mistral Large 4 beats Qwen 3.8 Max and Kimi K3 on Terminal-Bench — r/singularity

Get the daily brief of stories like this at 6:30 every morning →