AINewsnow

Qwen3.8-flash-next (iq2_xs)

strata model: qwen3.8-flash-next-iq2_xs engine: strata 0.1.33 context limit: 131,072 tokens gpu: rtx 4090, 24 gb system memory: 31 gb cpu: intel i9-14900k current settings - key/value cache: int8 - speculative decoding: enabled, 4-token window - expert cache: automatic - expert profile: enabled - e…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-02 19:21 · r/LocalLLM
    Qwen3.8-flash-next (iq2_xs)

More stories

  1. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
  2. Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
  3. Guided Vision in Gemini Live: built for accessibility — Google Gemini Blog
  4. Tavus unveils Griffin, the "first Human Interaction Model", which it says passed the "video Turing test", with 48% of users thinking it was human in live chats (@tavus) — Techmeme
  5. Google tests its plan for AI data centers in space with Project Suncatcher — Scientific American
  6. Google announces Gemini 4 Argon AI model, but you can't use it yet — Ars Technica AI
  7. OpenAI announces ‘dots’ agent after scrapping launch of new AI model over safety concerns — The Guardian AI
  8. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →