TensorFold vs vLLM on one DGX Spark, same benchmark: Qwen3.8-Flash-Next goes from 27.8 to 52.2 tok/s for a single request (1.4× with 5 at once)
I tested vLLM and TensorFold with Qwen3.8-Flash-Next on one DGX Spark (GB10, 128 GB unified memory). I used the same benchmark script and ran one server at a time. Setup is vLLM: nightly build, NVIDIA NVFP4 checkpoint, fp8 KV cache, MTP with 3 draft tokens, 6 slots x 262k context. TensorFold: v0.3.…
Read the full story at r/LocalLLM ↗