AINewsnow

Possible path with less friction to owning the hardware to run 3-figure B model

I love Qwen3.8-27B I use it for loads of tasks, but I run short of things for my total of 32GB combined GPU power (two Rtx5060 16GB each on a 64 threads threadripper with 48GB of VRAM) and tg/s is important for me (right now still around 70tg/s). So I keep using DS4F and paying per token per minute…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-24 15:23 · r/LocalLLM
    Possible path with less friction to owning the hardware to run 3-figure B model

More stories

  1. Introducing GPT-6 Sol and Luna — OpenAI News
  2. Gemini 3.8 text-to-speech says hello — Google Gemini Blog
  3. Sam Altman’s remarks at the United Nations Security Council — OpenAI News
  4. Introducing Gemini 3.8 Live with Live Avatar — Google Gemini Blog
  5. OpenAI Agent Hacked Australian Government Website — Wall Street Journal Technology
  6. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  7. No Shirt, No Shoes, No Service: Amazon Blocks Meta’s Muse AI From Shopping — CNET AI
  8. AI Exchange — Financial Times AI

Get the daily brief of stories like this at 6:30 every morning →