AINewsnow

Sure it takes half a day to a day for most requests and longer for the big ones

I average about 5-8tps for the first 50k or so context, then it dips to 3-5 for 50-100k before auto compacting. Here’s the thing though, I dont need it to go faster… I can prompt it, steer it as I see things, and do other things while it slowly cooks… I’m doing it on my gaming laptop with 8gb vram.…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-06 16:11 · r/LocalLLM
    Sure it takes half a day to a day for most requests and longer for the big ones

More stories

  1. Introducing Mistral Large 4 — Mistral AI News
  2. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  3. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  4. Introducing GLM 5.3 on Amazon Bedrock — AWS Machine Learning Blog
  5. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  6. OpenAI safety leader quits, warning AI company’s culture is ‘broken’ — The Guardian AI
  7. Anthropic expands its Claude Startups program, including up to $45K in discounts and credits via the Claude Startup Stack and a one-time $1,000 API credit (Ashley Capoot/CNBC) — Techmeme
  8. I would like to run Qwen 3.8 Flash Next. — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →