Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact
This story is from 2026-09-22. It is preserved in the archive; the latest stories are on the live feed.
Recently, the new DeepSeek-V4.1-Flash architecture showed how a causal encoder-decoder can work, but it was trained from scratch. Model Grafting does it to an existing model: cut at some depth, let the lower layers read the prompt, and use the upper layers get for encoder's residual stream as prefi…
Read the full story at r/huggingface ↗
Timeline · 1 report
- 2026-09-22 01:44 · r/huggingface
Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact