Open-source VLA and world-action inference runtime on Jetson Thor
We open-sourced InstinctFlash, an AGPL-3.0 inference runtime for VLA and world-action models. Demo: https://youtu.be/nku65iyL5Fw The runtime uses CUDA graphs, KV and conditioning-state caching, specialized attention paths, fused kernels, FP8 / mixed precision, and few-step diffusion distillation. O…
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-09-23 20:46 · r/deeplearning
Open-source VLA and world-action inference runtime on Jetson Thor