Used Opus 5.5 to optimize llama.cpp inference for Swift Qwen 3.8 27B Q6_K on RTX 5090 - decode 143 tok/s prefill 2840 tok/s
https://preview.redd.it/uflbt6x61tsh1.png?width=472&format=png&auto=webp&s=20499138331923a3d84505e7fa530d02d1a37978 https://preview.redd.it/toxiq4ah1tsh1.png?width=419&format=png&auto=webp&s=5cdcc3244accdeee500ddac8a5cac43e63c54851 https://preview.redd.it/21lkvj782tsh1.png?width=439&format=png&auto…
Read the full story at r/LocalLLM ↗