AINewsnow

I re-trained the DFlash 2 drafter for Ternary Bonsai 2 27B: 2.2x on an L4 (3.2x on code edits with ngram lookup), 1.5x on a Mac, 1.2x in Chrome

PrismML's Ternary Bonsai 2 27B fits a 24 GB card or Mac, but decodes at ~30 tok/s on an L4 and ~21 on an M4 Pro. z-lab's DFlash 2 drafter was trained on bf16 Qwen3.8-27B, so it guesses worse on the ternary model. I fine-tuned it on 1.5M tokens of Bonsai 2's own greedy output. Drafter (safetensors +…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-07 08:13 · r/LocalLLaMA
    I re-trained the DFlash 2 drafter for Ternary Bonsai 2 27B: 2.2x on an L4 (3.2x on code edits with ngram lookup), 1.5x on a Mac, 1.2x in Chrome

More stories

  1. EmbeddingGemma 2: an open, lightweight multimodal embedding model — Google DeepMind Blog
  2. Introducing Mistral Large 4 — Mistral AI News
  3. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  4. Sharing AI progress in mathematics — OpenAI News
  5. Google launches Playground, a browser-based, no-code AI game creation platform available to US users aged 18+, powered by Gemini, Nano Banana, and Lyria (Jay Peters/The Verge) — Techmeme
  6. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  7. Introducing the Decisions API — OpenAI YouTube
  8. Together Link: open models in the harness you already use. Start with one command today. — Together AI Blog

Get the daily brief of stories like this at 6:30 every morning →