AINewsnow

Qwen3.8-27B on 2x RTX 5070 Ti now matches an RTX 5090 in decode on the same weights (tensor-parallel NInfer fork)

I've been using a local Qwen3.8-27B as a sub-agent for Opus for a while now. For everyday coding it's more than good enough, and it saves a lot of tokens. For the harder, more novel stuff Opus is still needed. When I built my workstation I went with two RTX 5070 Ti instead of a 5090 to stay on budg…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-28 13:04 · r/LocalLLM
    Qwen3.8-27B on 2x RTX 5070 Ti now matches an RTX 5090 in decode on the same weights (tensor-parallel NInfer fork)

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. Launching Meta Enterprise Platform — Meta Newsroom
  3. Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents — The Guardian AI
  4. Scoop: Anthropic's Dario Amodei to have White House dinner with Trump — Axios AI+
  5. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  6. Bill Gates says unchecked AI could ‘cause a billion deaths’ in call for regulation — The Guardian AI
  7. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times Technology
  8. Nvidia’s Answer to Rogue Agents Is an Open-Source AI Security System — Wired AI

Get the daily brief of stories like this at 6:30 every morning →