AINewsnow

Small models are actually quite capable when the input is structured right...

We decided to run some small models, up to 3 billion parameters, on an old Note 8 (llama.cpp in Termux). We gave them what looked like a simple task: on the Wikipedia page for the Galaxy Note series, pick the "Note 8" link among similar ones ("Note 8.0", "Galaxy Note 8.0", "Note FE") and read the r…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-02 10:39 · r/LocalLLM
    Small models are actually quite capable when the input is structured right...

More stories

  1. TensorFold vs vLLM on one DGX Spark, same benchmark: Qwen3.8-Flash-Next goes from 27.8 to 52.2 tok/s for a single request (1.4× with 5 at once) — r/LocalLLM
  2. Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo — r/LocalLLM
  3. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers) — r/LocalLLM
  5. Browser FPS with 3D models, textures and SFX generated locally on one GPU, plus a local Qwen 27B for part of the code: my pipeline and what failed — r/LocalLLM
  6. Inside-Out AI: Rebuilding Airbnb Behind the Scenes and Across the Guest Experience — Latent Space
  7. Who’s the current “king” of local LLMs for you — Qwen, Gemma, Llama, something else? — r/LocalLLM
  8. Optimizing dual AMD r9700 setup — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →