New to LLM serving: how to share H200 VRAM across multiple models on Kubernetes?
Hi all, I’m new to serving LLMs and recently got access to an H200 rig. I’d like to get the most out of its VRAM, but I’m not sure of the right approach. Setup: GPUs: 8 × H200 (141 GB VRAM each) Kubernetes installed on the machine Goal: serve multiple LLMs plus some other AI models (e.g. [embedding…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-09 08:04 · r/LocalLLM
New to LLM serving: how to share H200 VRAM across multiple models on Kubernetes?
More stories
- GPT-6 and Intelligent UI for everyone — OpenAI News
- Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
- OpenAI Decisions API now available on AI Gateway — Vercel Blog
- Anthropic bans ‘abusive or cruel behavior’ toward Claude — The Verge AI
- Introducing Playground: Create and play custom games — Google AI Blog
- Introducing Mistral Large 4 — r/artificial
- Sophos cuts threat investigation time by 96% with OpenAI Daybreak — OpenAI News
- Grok Imagine Video 1.5 Lite on AI Gateway — Vercel Blog
Get the daily brief of stories like this at 6:30 every morning →