We cannot RDMA into a GPU's shared memory.
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
And that limitation turns out to explain why disaggregated inference is harder than the press releases suggest. A network can only write into one rung of any memory hierarchy: the one that's globally addressable. On a CPU that's DRAM. On a GPU that's HBM. Not L1, not SMEM, not tensor memory. NIXL (…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-08-20 23:53 · r/learnmachinelearning
We cannot RDMA into a GPU's shared memory.