AINewsnow

A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode

I forked Strata and worked with Opus 5.5 with some heavy changes to it to make it work on an IBM AC922 I have access to. The IBM AC922 is a 2018 era beast with two POWER9 20 core SMT4 CPUs that are connected by NVLink to 4 or 6 NVIDIA Tesla V100 SXM2 GPUs, the CPU-GPU BW advertised as 150GB/s and t…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-04 16:25 · r/LocalLLaMA
    A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode

More stories

  1. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
  2. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  3. Google launches satellite to test feasibility of building data centers in space — NPR Technology
  4. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  5. The 7-year-old Nvidia Shield TV is now $100 more expensive due to AI — Ars Technica AI
  6. NVIDIA Vera CPU Is Coming to CoreWeave: Pack In More Agents — CoreWeave Blog
  7. What It Takes to Bring Up a Multi-Rack NVIDIA Vera Rubin NVL72 Cluster — CoreWeave Blog
  8. What Comes Next: Operating and Evolving the Production AI Factory — CoreWeave Blog

Get the daily brief of stories like this at 6:30 every morning →