AINewsnow

TensorFold vs vLLM on one DGX Spark, same benchmark: Qwen3.8-Flash-Next goes from 27.8 to 52.2 tok/s for a single request (1.4× with 5 at once)

I tested vLLM and TensorFold with Qwen3.8-Flash-Next on one DGX Spark (GB10, 128 GB unified memory). I used the same benchmark script and ran one server at a time. Setup is vLLM: nightly build, NVIDIA NVFP4 checkpoint, fp8 KV cache, MTP with 3 draft tokens, 6 slots x 262k context. TensorFold: v0.3.…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-29 16:08 · r/LocalLLM
    TensorFold vs vLLM on one DGX Spark, same benchmark: Qwen3.8-Flash-Next goes from 27.8 to 52.2 tok/s for a single request (1.4× with 5 at once)

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  3. NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction — Hugging Face Blog
  4. Get Started with Kimi K3 on CoreWeave Dedicated Inference — CoreWeave Blog
  5. Add Runtime Controls to AI Agents with NVIDIA OpenShell — NVIDIA Technical Blog
  6. How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency — NVIDIA Technical Blog
  7. AMD Buys Fei-Fei Li's World Labs for $8.2 Billion to Challenge Nvidia — AlphaSignal
  8. Nvidia adds $150 billion to its historic stock buyback — Semafor Technology

Get the daily brief of stories like this at 6:30 every morning →