AINewsnow

Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38× faster decisions and +8.7 points accuracy, for under 2 GB extra memory

A few days ago I released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between options you define and returns a calibrated probability for each, in one forward pass. Speed was great on my M4 Max and RTX PRO 6000, but as a general zero-shot classifier it trailed the big models. Then it occ…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-10-01 13:59 · r/huggingface
    Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38× faster decisions and +8.7 points accuracy, for under 2 GB extra memory
  2. 2026-10-01 13:58 · r/LocalLLaMA
    Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38× faster decisions and +8.7 points accuracy, for under 2 GB extra memory

More stories

  1. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
  2. Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Allen Institute for AI (Ai2)
  3. Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
  4. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  5. Introducing dots — OpenAI News
  6. OpenAI Says It Will Not Release Newest Astra A.I. Model Over Safety Concerns — New York Times Technology
  7. Ollama now supports Jev-style decision models — Ollama Blog
  8. Introducing GPT-6.1 Sol — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →