AINewsnow

I optimized Bonsai 27B for 8 GB VRAM and agentic work: 36 tok/s

Ternary Bonsai 2 27B in its PTQ1_0 quant is 5.95 GB of weights, so all 65 layers stay on the card with a 64k window and the KV cache at q8_0/q4_0. I measured: RTX 4060 Ti 8 GB, CUDA: 36 tok/s AMD RX 570 8 GB, Vulkan/RADV: 7 tok/s The RX 570 is why I'm posting at all. Through the fork's Vulkan path…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-22 16:10 · r/LocalLLM
    I optimized Bonsai 27B for 8 GB VRAM and agentic work: 36 tok/s

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity — The Verge AI
  3. Alibaba Unveils New AI Chip, Outlines Plan for Larger Model — Wall Street Journal Technology
  4. Amazon blocks Meta’s Muse AI agent — The Verge AI
  5. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
  6. How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore — AWS Machine Learning Blog
  7. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  8. Systems for Machine Learning[D] — r/MachineLearning

Get the daily brief of stories like this at 6:30 every morning →