AINewsnow

Why the same Llama 3.2 1B model comes in different file sizes: a beginner’s explanation

I was confused by this, so i wrote a beginner friendly article explaining model quantization, what Q2, Q4, Q8, and F16 mean, and the trade-off between smaller files and precision. I have tried to explain it in the simplest way. Feedback or corrections are welcome.

Read the full story at r/learnmachinelearning ↗

Timeline · 1 report

  1. 2026-09-30 17:04 · r/learnmachinelearning
    Why the same Llama 3.2 1B model comes in different file sizes: a beginner’s explanation

More stories

  1. TensorFold vs vLLM on one DGX Spark, same benchmark: Qwen3.8-Flash-Next goes from 27.8 to 52.2 tok/s for a single request (1.4× with 5 at once) — r/LocalLLM
  2. Qwen flash next on 12+16gb vram, and 32gb ram viable? — r/LocalLLM
  3. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers) — r/LocalLLM
  5. Browser FPS with 3D models, textures and SFX generated locally on one GPU, plus a local Qwen 27B for part of the code: my pipeline and what failed — r/LocalLLM
  6. Who’s the current “king” of local LLMs for you — Qwen, Gemma, Llama, something else? — r/LocalLLM
  7. Optimizing dual AMD r9700 setup — r/LocalLLM
  8. I built an open-source tool that tells you why your vLLM server is slow (NVIDIA only for now, Mac support planned) — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →