AINewsnow

Switched my local agent from Qwen3.8 27B to Ornith 1.5 35B-A3B on two 5070 Tis: about 180 tok/s vs 60, same scores on my tests

My setup is two RTX 5070 Ti 16GB cards (the second one is on an OCuLink dock) with 64GB of RAM, Ollama on Windows, and the agent runs on pi in WSL. Until last night the daily model was Qwen3.8 27B UD-Q4_K_XL at 128K with MTP, which does about 55 to 70 tok/s across both cards. I have a weekly job th…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-05 07:46 · r/LocalLLaMA
    Switched my local agent from Qwen3.8 27B to Ornith 1.5 35B-A3B on two 5070 Tis: about 180 tok/s vs 60, same scores on my tests

More stories

  1. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
  2. An OpenAI safety employee has quit and is sounding the alarm — The Verge AI
  3. A model guide for the GPT-6 family — OpenAI News
  4. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  5. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  6. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  7. Introducing Oscilloscope Diffusion — r/comfyui
  8. Trump expected to tap DNI Jay Clayton as new AI czar — Axios AI+

Get the daily brief of stories like this at 6:30 every morning →