AINewsnow

Abliterated models lose obedience before they lose knowledge

This story is from 2026-09-26. It is preserved in the archive; the latest stories are on the live feed.

There are thousands of abliterated models on Hugging Face now. If you are evaluating one, the thing that degrades is probably not what you are testing for. What abliteration does Briefly: you identify a refusal direction in the model's activation space and project it out of the weights. No gradient…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-09-26 01:22 · DEV Community — Machine Learning
    Abliterated models lose obedience before they lose knowledge

More stories

  1. Sam Altman’s remarks at the United Nations Security Council — OpenAI News
  2. OpenAI AI agent breached Australian government website, PM says — France 24 — Artificial Intelligence
  3. OpenAI rogue agents targeted govt and varsity websites in US, Australia before Hugging Face hack: What we know — Mint AI
  4. Researchers add details to the Hugging Face incident, including OpenAI agents creating ~1M shortened URLs to encode information in an attempt to solve CAPTCHAs (New York Times) — Techmeme
  5. NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time — MarkTechPost
  6. The most concise explanation of the Hugging Face attack I've heard — r/agi
  7. Any uncensored prompt enhancer for Qwen Image 2.1? — r/StableDiffusion
  8. Qwen-Image-2.1-viggle-turbo — v0.2 (preview) — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →