I stored model-native numerical memory outside a frozen LLM and retrieved the associated memory through its own attention Qwen → Mistral replication, 127/128 Top-1
Two different 7B transformers. Two different internal coordinates. The same memory mechanism. I started AKBASCORE MAM on Qwen2.5-7B-Instruct. I have now independently localized and replicated the mechanism on Mistral-7B-Instruct-v0.3. Final Mistral result: 127/128 correct memories at Top-1 — 99.22%…
Read the full story at r/LocalLLM ↗