AINewsnow

I tried to write a C++ engine that makes Tensor-Train LLM layers run faster than dense FP16 on Apple Silicon (by using AMX utilization)

This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.

Everyone in the local LLM space uses INT4/INT8 quantization. It works perfectly for frozen models. But if you want to do on-device training or continuous learning, discrete quantization breaks gradient flow. Tensor-Train (TT) decomposition solves this by keeping the weights in a continuous Float32…

Read the full story at r/learnmachinelearning ↗

Timeline · 1 report

  1. 2026-08-22 09:16 · r/learnmachinelearning
    I tried to write a C++ engine that makes Tensor-Train LLM layers run faster than dense FP16 on Apple Silicon (by using AMX utilization)

More stories

  1. Meta's personal AI agent Muse climbs to No. 1 among free apps on Apple's US App Store, ahead of ChatGPT; Muse launched on September 8 (Georgia Hennessy/Business Insider) — Techmeme
  2. Meta's Muse AI agent is very powerful — but not enough to overcome my kids' annoying school apps — Business Insider AI
  3. Week in review: OpenAI ships managed Agents API, Apple's new Siri reportedly runs on Gemini, and three vendors add agent spend controls — r/artificial
  4. Apple reportedly building server packed with M-series Ultra chips for AI — Ars Technica AI
  5. Apple Reference Images Explained: The iPhone 18 Pro’s Hardware Solution to AI Slop — CNET AI
  6. Best open-source model for an M2 Max 32GB and what closed model does it actually compare to? — r/LocalLLM
  7. Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro — r/LocalLLaMA
  8. Apple M5 Ultra Scores Big GPU Gains in Leaked Geekbench Benchmark — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →