AI agents (mostly Opus 5.5) have been speeding up a 27B model from 66tok/s to 580tok/s on a Mac in 3 days by rewriting its inference engine
https://www.yukon.org/mlxfast
Read the full story at r/singularity ↗
https://www.yukon.org/mlxfast
Read the full story at r/singularity ↗
Get the daily brief of stories like this at 6:30 every morning →