AINewsnow

model : add Ling 3.0 VL support by aetherbird · Pull Request #29151 · ggml-org/llama.cpp

Model Overview Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-09-24 13:37 · r/LocalLLaMA
    vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (… · ggml-org/llama.cpp@70c4e15
  2. 2026-09-24 08:36 · r/LocalLLaMA
    model : add Ling 3.0 VL support by aetherbird · Pull Request #29151 · ggml-org/llama.cpp

More stories

  1. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLaMA
  2. Jev) Mica 4B vs Laya on Tetris: same seed, same prompt, 0 output tokens, running locally on llama.cpp — r/LocalLLM
  3. Qwen3-Coder 30B on RTX 3080 20GB — KV cache stability and context tuning — r/LocalLLM
  4. PSA: llama.cpp -cram should be increased for agentic workflows (default is 8192) — r/LocalLLaMA
  5. Gufo: the all-in-one strix halo inference engine — r/LocalLLM
  6. My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context — r/LocalLLaMA
  7. I turned Qwen3.8-27B Q2_64 + llama.cpp into a fully TypeSafe AI-compatible Jev-like system. OpenAI API still intact! World’s first Vision-enabled Jev-like model! <10 GB VRAM, 170 ms on an RTX 3090 and ~140 tok/s in chat. 76% vs. 88% Jev-1.13 Acc. on a diverse 22,000-request typed-decision benchmark — r/LocalLLaMA
  8. GGUFs in transformers natively! — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →