AINewsnow

I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window.

**DISCLAIMER** THE PREFILLING TPS SHOWN ON THE PHONE IS COMPUTED ONLY FOR THE LAYERS IT HOLDS. ALREADY FIXING IT TO SHOW END-TO-END PREFILL RATE. NUMBERS BELOW ARE ACCURATE FOR E2E PREFILL RATE. Every file or tool result my agent reads on a 24 GB M4 Pro MacBook is a wait, and 64k of 8-bit context i…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-10-02 17:16 · r/LocalLLM
    I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window.
  2. 2026-10-02 16:59 · r/LocalLLaMA
    I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window.

More stories

  1. Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune) — r/LocalLLM
  2. Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo — r/LocalLLM
  3. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Some topics are off limits in Chinese AI, researchers find — CBS News Technology
  5. What is your experience with bonsai 2 27b? — r/ArtificialInteligence
  6. Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print. — r/LocalLLaMA
  7. Browser FPS with 3D models, textures and SFX generated locally on one GPU, plus a local Qwen 27B for part of the code: my pipeline and what failed — r/LocalLLM
  8. Continuity update: screen replacement, image to 3D, Qwen Image 2.1, and a lot more since 3.0 — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →