AINewsnow

Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity

I’ve been experimenting with whether Qwen3.8-Flash-Next’s pretrained PLE n-gram memory can improve a much smaller Qwen3.5-0.8B model. I trained the 0.8B setup with limited resources, mostly using free Kaggle notebook GPUs. The setup keeps both the Qwen3.5-0.8B backbone and the roughly 51B-parameter…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-25 12:46 · r/LocalLLaMA
    Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity

More stories

  1. Need some help — r/AI_Agents
  2. DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes) — r/LocalLLaMA
  3. Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation — MarkTechPost
  4. Perplexity's Photon Slashes Search Latency by 12x for AI Agents — AlphaSignal
  5. asked gemini and chatgpt the same question about local businesses and they recommended almost completely different ones. only 11% of the sources they cite overlap — r/GeminiAI
  6. open an incognito tab and ask chatgpt for the best [what you do] in your city. most owners have never checked whether they come up, and the numbers are worse than you'd think — r/PromptEngineering
  7. Bros benchmark is logo design — r/ArtificialInteligence
  8. Introducing GPT-6 Sol and Luna — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →