Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.
Read the full story at AWS Machine Learning Blog ↗
Timeline · 2 reports
- 2026-09-10 21:37 · AWS Machine Learning Blog
Reduce inference cold starts on Amazon SageMaker HyperPod with model caching - 2026-09-09 22:26 · AWS Machine Learning Blog
Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM