$0 compute, 5 architectures, 16 runs: surgical data poisoning makes LLMs indifferent [margin -> 0.0] while PPL looks fine. I built a 0.1ms gate that stops it
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
TL;DR: Fine-tuned 5 open LLMs on a stream with 50-70% lies. Without defense, truth margin collapses to ~0.0 - the model becomes indifferent between truth and lie - while PPL looks healthy. Built Beatriz, a non-invasive proxy gate. Gate alone gives 65% of benefit without touching the student loop. F…
Read the full story at r/machinelearningnews ↗
Timeline · 1 report
- 2026-09-09 17:57 · r/machinelearningnews
$0 compute, 5 architectures, 16 runs: surgical data poisoning makes LLMs indifferent [margin -> 0.0] while PPL looks fine. I built a 0.1ms gate that stops it