A 0.8B model just beat a 2B model on ARC-Challenge (42.15%): Closed-form weight surgery beat multi-GPU SFT with 0 backprop (Independently verified on NVIDIA L4)
A few days ago we shared the idea behind DynamicTune: transferring the trajectory flow from a larger teacher model directly into a smaller student via closed-form linear algebra in ~12 minutes on consumer hardware. Zero backpropagation, zero training tokens, zero gradient descent. To eliminate loca…
Read the full story at r/LocalLLaMA ↗