One post-training framework from 24 GB QLoRA to multi-node SFT and GRPO
i work on Halo at White Circle. we built Halo so moving from a small LoRA experiment to a distributed run does not require replacing the model, trainer or checkpoint format. the repo includes single-GPU LoRA and QLoRA examples, plus SFT, DPO, SMPO, online GRPO, environment GRPO, distillation and re…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-21 17:47 · r/LocalLLM
One post-training framework from 24 GB QLoRA to multi-node SFT and GRPO