Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact
Recently, the new DeepSeek-V4.1-Flash architecture showed how a causal encoder-decoder can work, but it was trained from scratch. Model Grafting does it to an existing model: cut at some depth, let the lower layers read the prompt, and use the upper layers get for encoder's residual stream as prefi…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-22 01:44 · r/huggingface
Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact - 2026-09-22 01:43 · r/LocalLLaMA
Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact