Memory-Efficient Pre-training of a 1.11B LLM on a 6 GB Consumer GPU
​ BAdam + CPU offload + BitNet ternary weights + tied embeddings + chunked cross-entropy. DOI: 10.5281/zenodo.22725431 Full write-up: https://huggingface.co/blog/RitishReal/1b-petrain-llm-in-4gb-vram
Read the full story at r/machinelearningnews ↗
Timeline · 1 report
- 2026-09-26 15:29 · r/machinelearningnews
Memory-Efficient Pre-training of a 1.11B LLM on a 6 GB Consumer GPU