AINewsnow

We swapped AdamW's optimizer states for a Fast Fourier Transform (FFT) to cut VRAM in half. Anyone else trying non-quantization methods?

Hey everyone, Like most of you, we have been fighting constant OOM errors while trying to fine-tune 8B and 70B models on consumer GPUs. The AdamW optimizer states are always the biggest bottleneck. We didn't want to rely on aggressive 8-bit quantization because we were seeing degradation in converg…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-05 19:54 · r/LocalLLaMA
    We swapped AdamW's optimizer states for a Fast Fourier Transform (FFT) to cut VRAM in half. Anyone else trying non-quantization methods?

More stories

  1. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  2. OpenAI safety employee resigns, claiming the company’s ‘culture is broken’ — TechCrunch AI
  3. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  4. Nvidia-Backed Reflection Unveils Open AI Model, Taking on China — Bloomberg AI
  5. Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen — r/LocalLLaMA
  6. To comply with the EU AI Act, OpenAI plans to add text watermarking for ChatGPT and Codex users in the EU and an opt-in setting for API customers globally (OpenAI) — Techmeme
  7. Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 — r/LocalLLaMA
  8. Supercharge regulated workloads with Claude Code and Amazon Bedrock — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →