AINewsnow

PyTorch 2.14 Ships Faster Apple Silicon, Rebuilt Distributed Training and NVGEMM

This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.

The latest PyTorch release ships NVGEMM CUTLASS kernels for Inductor, a rebuilt nccl2 backend, native Apple Silicon linear algebra, and first-class fault tolerance.

Read the full story at AlphaSignal ↗

Timeline · 1 report

  1. 2026-09-02 18:56 · AlphaSignal
    PyTorch 2.14 Ships Faster Apple Silicon, Rebuilt Distributed Training and NVGEMM

More stories

  1. [Release] Nirvana Code: A single-binary Rust LLM engine built from the metal up for Apple Silicon (Metal 3, Persistent Prefix Cache, Speculative Decoding, Dual GGUF + MLX) — r/LocalLLM
  2. He’s the Face of AI Doomsday Fears — Wall Street Journal Technology
  3. Week in review: OpenAI ships managed Agents API, Apple's new Siri reportedly runs on Gemini, and three vendors add agent spend controls — r/artificial
  4. Running ACE-Step 1.5 and YuE2-3B on one GPU behind a single local UI (CUDA + Apple Silicon): notes from building it — r/LocalLLM
  5. Best open-source model for an M2 Max 32GB and what closed model does it actually compare to? — r/LocalLLM
  6. Apple M6 Pro Geekbench 7 — r/LocalLLaMA
  7. Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro — r/LocalLLaMA
  8. Apple M5 Ultra Scores Big GPU Gains in Leaked Geekbench Benchmark — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →