Mac and Qwen 3.8 27B users... are you using GGUF or MLX? I need 100K of context and 10-15 tps with Q4.
This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.
Hi people. Someone recommended to me using gguf instead of mlx because mlx would consume too much memory on larger context. And to use llamacpp directly. I have 32gb of ram btw. What is your recommendation? I didn't get great results with LMStudio (limits my context too much) and was trying MTPLX b…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-03 13:38 · r/LocalLLM
Mac and Qwen 3.8 27B users... are you using GGUF or MLX? I need 100K of context and 10-15 tps with Q4.