AINewsnow

Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact

Recently, the new DeepSeek-V4.1-Flash architecture showed how a causal encoder-decoder can work, but it was trained from scratch. Model Grafting does it to an existing model: cut at some depth, let the lower layers read the prompt, and use the upper layers get for encoder's residual stream as prefi…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-09-22 01:44 · r/huggingface
    Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact
  2. 2026-09-22 01:43 · r/LocalLLaMA
    Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact

More stories

  1. Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
  2. XiaomiMiMo/MiMo-V2.6-Pro-RL · Hugging Face — r/LocalLLaMA
  3. Jev's calibration was measured. The LLMs won [D] — r/MachineLearning
  4. DeepSeek's Liang: Next Models Must Train on Huawei and Domestic Chips — Pandaily
  5. Free GLM 5.3 Flash and Deepseek V4.1 for a month on 100% Private US Architecture — r/ChatGPTCoding
  6. Local harness test: omp vs qwen code vs deepseek harness (fresh install) — r/LocalLLM
  7. next-skill-router: cost-aware, local-first skill routing with Ollama/vLLM — r/machinelearningnews
  8. Deepseek training 2T and plans 8T model — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →