AINewsnow

Is it truly impossible to stream the model from SSD to RAM with a decent token per second? Is it truly no possible hardware, software or even architectural solution?

This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.

As far as I understand LLM's only multiply one parameter at a time instead of using the whole model at once so is it truly impossible to do efficient streaming from SSD. In video games the whole map is not loaded at once only the part that you are in is loaded and streaming is soo efficient there y…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-02 10:06 · r/LocalLLM
    Is it truly impossible to stream the model from SSD to RAM with a decent token per second? Is it truly no possible hardware, software or even architectural solution?

More stories

  1. Anthropic says Claude 'leads' 26 percent of its AI R&D work — Engadget
  2. Google announces new experimental "CC" AI agent for families — Ars Technica AI
  3. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  4. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  5. Optimizing agent system prompts with Amazon Bedrock AgentCore — AWS Machine Learning Blog
  6. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  7. Introducing Astra for Law — OpenAI News
  8. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →