AINewsnow

Ternary Bonsai 2 (27B) fails to load in LM Studio and oMLX. I made fixes for both (GGUF PQ2_0/PTQ1_0 + MLX 2-bit)

Prism ML's Ternary-Bonsai-2-27B is a Qwen3.8-27B at ~1.7 bits/weight (5.9–7.2 GB GGUF, 8.4 GB MLX). But the popular runners don't load it yet: LM Studio: invalid ggml type 142. should be in [0, 43) . PQ2_0 / PTQ1_0 are custom ternary types that only exist in Prism's llama.cpp fork. oMLX: Model type…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-21 08:01 · r/LocalLLM
    Ternary Bonsai 2 (27B) fails to load in LM Studio and oMLX. I made fixes for both (GGUF PQ2_0/PTQ1_0 + MLX 2-bit)

More stories

  1. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  2. Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF) — r/LocalLLM
  3. Qwen3.8-27B on a single RTX 5090 (32GB) + 64GB DDR5-6000 — looking for real t/s numbers (llama.cpp / vLLM / sglang) — r/LocalLLM
  4. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  5. Yandex open-sourced an 80B model trained from scratch: what's inside and where it wins — DEV Community — Machine Learning
  6. You can use any LLM just like JEV — r/LocalLLaMA
  7. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  8. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →