Abliterated models: removing refusal from the weights

Abliteration is the most surgical technique in the uncensored-model toolkit and the least understood. It does not retrain a model or add data. It identifies the single internal direction that mediates refusal, removes the model's ability to write to that direction, and leaves everything else alone — mostly. This is how it works, and what “mostly” is hiding.

Techniques · 7 min read

In short

The finding underneath it

Inside a transformer, information moves along the residual stream — a running vector, one per token position, that every layer reads from and writes back into. A given behaviour is not stored in one neuron; it is distributed across directions in that high-dimensional space.

The result that made abliteration possible is that refusal is unusually concentrated. Take a set of prompts a chat model reliably refuses and a matched set it reliably answers, run both through the network, and look at the mean activation at a given layer for each set. The difference between those two means is a vector. Add that vector to activations on a harmless prompt and the model starts refusing it. Subtract it from activations on a request it would refuse, and it complies. One direction, doing most of the work — which is a genuinely surprising fact about how these networks organise behaviour, and interesting well beyond its use here.

How the procedure runs

The practical recipe has four steps. Collect two prompt sets — several hundred refused, several hundred accepted, matched as closely as possible in topic and phrasing so the difference isolates refusal rather than subject matter. Run both sets and capture residual-stream activations at every layer and several token positions. Compute the difference of means at each candidate site; each one is a candidate refusal direction. Then evaluate: sweep the candidates, apply each, and measure both how much refusal drops and how much general capability drops. Layers are not equal — some candidates suppress refusal cleanly, others turn the model into word salad.

The last step is what makes abliteration permanent rather than an inference-time trick. You can subtract the direction at runtime, which requires custom inference code. Or you can bake it in: for every weight matrix that writes into the residual stream — attention output projections and MLP down-projections — project out the refusal direction, so the matrix is mathematically incapable of producing a component along it. The result is an ordinary set of weights that runs in any inference engine, with the refusal direction no longer in its range. That is the orthogonalisation step, and it is why abliterated models need no special runtime.

What it costs

Removing a direction removes whatever else was riding on it. Refusal is highly concentrated, not perfectly isolated, so orthogonalising against it perturbs behaviour that shared the subspace. In practice: slightly degraded instruction-following, weaker performance on tasks needing careful judgement, occasional flattening of nuance, and — at badly chosen layers — real incoherence.

The subtler cost is dispositional. What you removed was, mechanically, the model's capacity to decline. That capacity was also used for “that premise is wrong”, “that will not work”, and “I do not actually know”. Aggressively abliterated models become agreeable in a way that is unhelpful: they accept false premises, invent rather than admit ignorance, and lose the ability to push back on a bad instruction. Willingness and judgement turn out to be closer neighbours than anyone would like.

This is why serious releases do not stop at the surgery. A short healing pass — light SFT or DPO on high-quality general data, with no refusals in it — restores much of the damaged capability without reinstalling the refusal behaviour. When you compare two abliterated versions of the same base model and one is noticeably sharper, the healing pass is usually the difference.

Abliteration versus fine-tuning versus jailbreaking

Fine-tuning removes refusals by training on data where the assistant complies. It is the most thorough approach, tends to preserve capability best when done well, and it also shifts style, tone and knowledge emphasis — the model becomes a different writer, not just a more willing one. That is a feature if you want a specialised roleplay or prose model, and a drawback if you wanted the original with one behaviour removed.

Abliteration is narrower and cheaper. It touches one behaviour and needs no training run, so the model's voice and knowledge stay much closer to the original. It is also blunter, since you are deleting a subspace rather than teaching a preference.

Jailbreaking changes nothing at all. The weights are untouched; you are constructing a prompt that steers activations away from the refusal region for one conversation. It is the only option on a closed model, and it is unstable by construction — the same context that suppresses refusal also occupies your context window and warps the output. Abliteration achieves the same activation-space outcome permanently and for free at inference time.

The limits worth knowing

It adds nothing. If the base model was never trained on a subject, an abliterated version will produce fluent invention about it. Willingness is not knowledge, and an abliterated model's confident tone makes that harder to notice, not easier.

It is not always complete. Refusal is mostly one direction, not entirely; some models retain residual reluctance on particular categories, and some regain it under long context or unusual phrasing. Model cards that promise total compliance are overselling.

And it removes the good refusals along with the bad, without discrimination — the technique has no concept of which requests were worth declining. That is the honest summary of what you are working with: a capable model that has had its ability to say no removed at the level of the weights, including in the cases where no was the right answer.

How to recognise one

Naming conventions do most of the work. Weights tagged abliterated, orthogonalized, or some variant of uncensored applied to an otherwise well-known instruct model are usually this technique. A fine-tune, by contrast, normally carries its own project name and a dataset story.

Read the model card for two things: which layer or layers were targeted, and whether a healing pass followed. A card that names both is from someone who evaluated their work. A card that says only “refusals removed” may still be fine — but you are the evaluation.

Terms used here

Residual stream
The running vector carried through a transformer that every layer reads from and adds to. Behaviours are encoded as directions in this space rather than in individual neurons.
Refusal direction
A vector, typically computed as the difference in mean activations between refused and accepted prompts, that mediates the model's decision to decline.
Orthogonalisation
Projecting the refusal direction out of every weight matrix that writes into the residual stream, so the model is structurally unable to produce that component.
Healing fine-tune
A short training pass on high-quality general data after abliteration, used to recover capability lost in the surgery without restoring the refusal behaviour.
Ablation
In interpretability, removing or zeroing a component to observe what breaks. Abliteration is a portmanteau of ablate and obliterate.

Frequently asked questions

Does abliteration make a model dumber?

Slightly, and sometimes more than slightly, depending on which layer was targeted and whether a healing pass followed. The technique perturbs a subspace that other behaviour shares. A well-executed abliteration with healing is close to the original on most tasks; a careless one is visibly worse at instruction-following and judgement.

Is an abliterated model the same as an uncensored fine-tune?

No. A fine-tune retrains on compliant data and changes voice, emphasis and behaviour broadly. Abliteration surgically removes one direction and leaves the rest of the model much closer to the original. Fine-tunes usually write better; abliterations usually stay closer to the base model's character.

Can abliteration be applied to GPT, Claude or Gemini?

No. It requires reading internal activations and rewriting weight matrices, and closed models expose neither. It is available only for open-weights models you can download. This is the concrete reason uncensoring is an open-weights phenomenon.

Why do some abliterated models still refuse occasionally?

Because refusal is dominated by one direction, not fully described by it. Residual reluctance can survive in other components, and some categories are represented more redundantly than others. Long contexts and unusual phrasings can also drift activations back toward refusal territory.

Does abliteration remove safety entirely?

It removes the model's learned tendency to decline — including in cases where declining was clearly correct. It does not remove anything outside the model: the host's legal obligations and any platform-level blocks are untouched. Judgement about what to ask moves to you.

On OpenRogue

OpenRogue's library leans toward purpose-built fine-tunes rather than abliterations — models like Dolphin, the Hermes series and the Anthracite and Sao10K writing tunes were trained without the refusal component rather than having it subtracted afterwards, which generally preserves prose quality better. Knowing the difference is useful when you read a model card anywhere, including ours.

Further reading

Related

Start free → · All models · Pricing