AINewsnow

We just open-sourced the world's fastest WebGPU kernels for local AI on Hugging Face

The collection includes kernels for more than 200 common ML operations, all of which can run entirely locally in your browser on WebGPU. We're also working to upstream these optimizations to Transformers.js, ONNX Runtime Web, LiteRT.js, and more! Kernels: https://huggingface.co/kernels?platform=web…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-30 16:02 · r/LocalLLaMA
    We just open-sourced the world's fastest WebGPU kernels for local AI on Hugging Face

More stories

  1. Cloudflare debuts open-weight multimodal decision models Clef and Clef-flash, claiming they are smarter and faster than Jev, based on Qwen3.8-27B and Qwen3.5-9B (Brandon Vigliarolo/The Register) — Techmeme
  2. OpenAI hit with landmark lawsuit following Hugging Face hack — Axios AI+
  3. Qwen flash next on 12+16gb vram, and 32gb ram viable? — r/LocalLLM
  4. Rogue AI agents: A timeline of security breaches since the attack on Hugging Face — Fast Company AI
  5. Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence — MarkTechPost
  6. FTC is investigating OpenAI, Anthropic and other AI companies over product risks — CNBC Technology
  7. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
  8. Frog and Toad and the Increasingly Capable Machines — LessWrong (Curated)

Get the daily brief of stories like this at 6:30 every morning →