AINewsnow

PSA: if you're on an Intel hybrid CPU, run Strata's calibrate - it nearly tripled my decode speed (IQ3_S at 256K, 16 GB card)

This is the third version. In the third version, I added P-core thread pinning. setup: 5070 ti 16gb, 96gb ram, i7-14700kf, windows. qwen3.8-flash-next iq3_s on strata, 262k context. first test: ~17 tok/s at 256k. log screenshot attached, before lines are stock settings. then i changed 3 things: poo…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-04 19:11 · DEV Community — Machine Learning
    PSA: if you're on an Intel hybrid CPU, run Strata's calibrate - it nearly tripled my decode speed (IQ3_S at 256K, 16 GB card)

More stories

  1. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
  2. An OpenAI safety employee has quit and is sounding the alarm — The Verge AI
  3. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  4. A model guide for the GPT-6 family — OpenAI News
  5. OpenAI fires 3 AI safety researchers for allegedly sharing confidential company information — Mint AI
  6. Introducing Oscilloscope Diffusion — r/comfyui
  7. Apple says it's tightening macOS Full Disk Access' controls due to new risks from AI agents — TechCrunch AI
  8. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →