Prefil of local models vs opus and astra
Why does no one talk about what the prefil speeds of these API providers are vs running locally. People with sparks or strix halos only seem to focus on decode without considering how much slower it is because of slow pp. Are there any benchmarks / figures of how fast the APIs process input? submit…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-19 05:50 · r/LocalLLaMA
Prefil of local models vs opus and astra