M2 Mac ultra128gb Qwen flash next
I am trying to get good speeds for Qwen flash next but I also need some ram for the system. Right now gpu is at 96gb around more I cannot use. Before I was getting 50 token/s from omlx and 60 tokens/s from mlx serve but lower context like 32k. I had ngram on ssd now where my system has more ram for…
Lead source: r/LocalLLM