GigaChat-3.5-Reasoning
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
Hey y'all! We've released a new model in our lineup: GigaChat-3.5 Reasoning. It's a 432B-A28B MoE with Gated DeltaNet for long-context efficiency. We trained domain experts (code, math, general, etc.) with CISPO and then distilled them into a single model via on-policy distillation. In our evals th…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-10 12:19 · r/LocalLLaMA
GigaChat-3.5-Reasoning