AINewsnow

Custom Node for INT8 ConvRot on Apple Silicon GPU

This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.

PyTorch has no aten::_int_mm kernel for the MPS backend. Every int8 linear layer therefore round-trips to the CPU, which turns an int8 model from the fastest thing you can run on a Mac into the slowest. This patch computes those matmuls on the GPU instead.

Read the full story at r/comfyui ↗

Timeline · 1 report

  1. 2026-09-03 16:06 · r/comfyui
    Custom Node for INT8 ConvRot on Apple Silicon GPU

More stories

  1. [Release] Nirvana Code: A single-binary Rust LLM engine built from the metal up for Apple Silicon (Metal 3, Persistent Prefix Cache, Speculative Decoding, Dual GGUF + MLX) — r/LocalLLM
  2. He’s the Face of AI Doomsday Fears — Wall Street Journal Technology
  3. Week in review: OpenAI ships managed Agents API, Apple's new Siri reportedly runs on Gemini, and three vendors add agent spend controls — r/artificial
  4. Running ACE-Step 1.5 and YuE2-3B on one GPU behind a single local UI (CUDA + Apple Silicon): notes from building it — r/LocalLLM
  5. Best open-source model for an M2 Max 32GB and what closed model does it actually compare to? — r/LocalLLM
  6. Apple M6 Pro Geekbench 7 — r/LocalLLaMA
  7. Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro — r/LocalLLaMA
  8. Apple M5 Ultra Scores Big GPU Gains in Leaked Geekbench Benchmark — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →