AINewsnow

Qwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW

Someone recommended that I try the ByteShape Qwen 3.8 27B IQ3-XXS GGUF after seeing my previous testing of the GSQ quant. So I did. And the result was… surprisingly bad. For context, I'm running: RTX 3060 12GB 16GB DDR4 RAM, single channel CachyOS / Arch Linux llama.cpp Qwen 3.8 27B MTP/speculative…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-22 16:31 · r/LocalLLaMA
    Qwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW

More stories

  1. yandex/AliceAI-Foundation-80B-A3B-Base: Russian-developed competitor to Qwen 35B and DeepSeek V4 Flash — r/LocalLLaMA
  2. Dual B60 24GB Performance — r/LocalLLM
  3. My contribution to the local AI community: 9 abliterated models, 99 GGUF quantizations in progress — r/huggingface
  4. Is llama.cpp meant to be slow at long context, even when you aren't using that context? — r/LocalLLaMA
  5. How I structured 50k synthetic ICD-10 QA pairs for local LLM fine-tuning — r/deeplearning
  6. Who's getting above 50 tok/s on AMD 9070, R9700 GPUs? — r/LocalLLM
  7. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  8. Help? — r/GeminiAI

Get the daily brief of stories like this at 6:30 every morning →