AINewsnow

I run a persistent local agent on a 16GB Air. Qwen3-8B on the GPU, a second brain on the Neural Engine, no cloud.

Her name is Bad Apple. Rust core, Swift platform layer, everything on the machine. She sees, remembers, audits everything she does, and comes back after restarts. Repo: https://github.com/savageAZfck/Bad_Apple ## The brains - **GPU:** Qwen3-8B-4bit on MLX. Chat, tools, verification. ~10 tok/s decod…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-09 23:20 · r/LocalLLM
    I run a persistent local agent on a 16GB Air. Qwen3-8B on the GPU, a second brain on the Neural Engine, no cloud.

More stories

  1. I built an open source framework for building RL environments. Named "Seahaven" after the fake town in The Truman Show. — r/LocalLLaMA
  2. Nara Baby introduced paid plans, so I tried building our own baby tracker with Claude Code — r/ClaudeAI
  3. Claude created my dream game, and got approved for Apple iOS store! — r/ClaudeAI
  4. Whallm 1.1.11: Swift1.5-Qwen3.8 support, and Qwen3.8 now runs at 15–17 tok/s decode and 500–600 tok/s prefill at 16K on a 64 GB Mac — r/LocalLLM
  5. Switch Billing - Lose Resets? — r/OpenAI
  6. What would make you actually use a personal AI assistant everyday? — r/artificial
  7. My ComfyUI Nodes and Workflows - Krea 2 (Turbo and Raw), Z-Image (Turbo and Base), MiniMax Music 3, Image2Text and LLM Chat (with Tools), Torch, Apple MLX and Cloud — r/comfyui
  8. MTP in llama.cpp now decodes competitively with ds4 using GLM 5.3 Flash — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →