[P] FP8 on the wire for VLA training: what actually paid off on a PCIe-only 8-GPU box, and the 48%-fewer-bytes change that bought exactly 0 ms
We train a three-modality VLA model (video diffusion + action expert + a frozen VLM for understanding, mixture-of-transformers style). Originally on 8×A800 with NVLink; we're moving it to 8× workstation-class Blackwell cards, which means no NVLink and no NVSwitch — every GPU-to-GPU byte goes over P…
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-09-20 13:28 · r/deeplearning
[P] FP8 on the wire for VLA training: what actually paid off on a PCIe-only 8-GPU box, and the 48%-fewer-bytes change that bought exactly 0 ms
More stories
- Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
- Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
- NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
- AI hallucination of Chinese nuclear components almost led to US military attack — Ars Technica AI
- TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text — MarkTechPost
- Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use — r/machinelearningnews
- Claude, Anthropic’s AI model, is helping to develop the next version of itself — Fast Company AI
- Amazon SageMaker Inference: 2026 year-to-date launches in review — AWS Machine Learning Blog
Get the daily brief of stories like this at 6:30 every morning →