I built an open-source tool that tells you why your vLLM server is slow (NVIDIA only for now, Mac support planned)
Heads up for Mac folks first: this currently works only on NVIDIA GPUs with vLLM. llama.cpp, MLX and Apple Silicon support are on the roadmap, and SGLang is too. Tensward runs your existing vLLM setup against your own prompts, reads the engine's metrics, compares the results with the GPU's theoreti…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-01 17:51 · r/LocalLLM
I built an open-source tool that tells you why your vLLM server is slow (NVIDIA only for now, Mac support planned)