AINewsnow

SplitMOE: experts with shared dimensions [ suprisingly worked near or better than standard MOE architecture ]

This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.

The intuition is similar to the DeepSeek MOE, but architecture is different. Please do check it out, and let me know your views. [ https://github.com/Priyanshu-5257/SplitMoE ]

Read the full story at r/deeplearning ↗

Timeline · 1 report

  1. 2026-09-11 07:13 · r/deeplearning
    SplitMOE: experts with shared dimensions [ suprisingly worked near or better than standard MOE architecture ]

More stories

  1. The Sequence Learning Loop - Issue 934: Understanding DeepSeek V4.1 Flash, DeepMind’s AlphaGenome Atlas and Muse — TheSequence
  2. DeepSeek’s Insane New Architecture — Two Minute Papers
  3. Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs — MarkTechPost
  4. Gemini 4 is good enough - JUST RELEASE IT — r/GeminiAI
  5. Gemini 4 หลุด แต่ paper ที่ Google เพิ่งตีพิมพ์ตรวจสอบได้ทุกตัวเลข — DEV Community — AI
  6. Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI
  7. I enjoyed the daily HF papers today — r/LocalLLaMA
  8. Engrams Embedding Entendre: Codesign for Efficient DRAM/SSD Offloading — SemiAnalysis

Get the daily brief of stories like this at 6:30 every morning →