AINewsnow

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.

Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.

Read the full story at AWS Machine Learning Blog ↗

Timeline · 2 reports

  1. 2026-09-10 21:37 · AWS Machine Learning Blog
    Reduce inference cold starts on Amazon SageMaker HyperPod with model caching
  2. 2026-09-09 22:26 · AWS Machine Learning Blog
    Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

More stories

  1. Can ChatGPT be an alternative to subscription software and can it really cut costs? — r/ChatGPTCoding
  2. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  3. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  4. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  5. Google's Gemini AI hacks three other companies during security test — Sky News Technology
  6. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  7. OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system — The Guardian AI
  8. A shared agentic platform for Wood Mackenzie, on Amazon Bedrock AgentCore — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →