AINewsnow

DeepSeek-V4-Pro distilled into Qwen3.5 9B/4B — a pretty interesting pair of compact reasoning models

This story is from 2026-08-21. It is preserved in the archive; the latest stories are on the live feed.

If you're looking for smaller Qwen3.5-based reasoning models that can still run locally, these two DeepSeek-V4-Pro distilled checkpoints are worth checking out: DeepSeek-V4-Pro-Qwen3.5-9B DeepSeek-V4-Pro-Qwen3.5-4B Both use the Qwen3.5 9B/4B family as the student models and transfer reasoning patte…

Read the full story at r/huggingface ↗

Timeline · 1 report

  1. 2026-08-21 14:18 · r/huggingface
    DeepSeek-V4-Pro distilled into Qwen3.5 9B/4B — a pretty interesting pair of compact reasoning models

More stories

  1. Bolt Adds DeepSeek V4.1 Flash at 10x Cheaper Than V4 Pro — AlphaSignal
  2. Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs — MarkTechPost
  3. Gemini 4 is good enough - JUST RELEASE IT — r/GeminiAI
  4. A 2026 Guide to Multi-Model AI Apps: GPT, Claude, Gemini & DeepSeek — DEV Community — AI
  5. Is there any use of a local llm with a 20B LLM? — r/AI_Agents
  6. Ai used for chatbots — r/artificial
  7. I gave 6 different AIs the same 5 questions — r/AI_Agents
  8. PromptDeck v1.1.0 – open-source desktop app to benchmark local AND cloud LLMs side-by-side (Ollama, LM Studio + OpenRouter, Groq, DeepSeek…) — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →