AI alignment and why models refuse

Nobody wrote a list of forbidden questions. Refusals emerge from a training process that rewards caution and cannot tell the difference between a dangerous request and a request that superficially resembles one. Understanding that process explains almost every frustrating thing a mainstream assistant does — and explains why the fix has to happen in the weights.

Alignment · 8 min read

In short

The pipeline that produces a refusal

A chat model is built in stages. Pretraining does the heavy lifting: predict the next token across an enormous corpus until the network has absorbed grammar, facts and reasoning patterns. The result is a base model — fluent, knowledgeable, and useless as an assistant, because it will happily continue your question with three more questions instead of answering it.

Supervised fine-tuning comes next. Show the model tens of thousands of well-formed instruction/response pairs and it learns the shape of an assistant turn. Some refusal behaviour enters here simply because the demonstration set contains refusals.

Then comes preference optimisation, which is where refusals are really installed. In classic RLHF, humans rank pairs of candidate responses, a reward model is trained to predict those rankings, and the policy is updated to maximise predicted reward — historically with PPO. Direct Preference Optimization skips the separate reward model and optimises the policy against the preference pairs directly, which is cheaper and now extremely common. Constitutional or AI-feedback approaches replace some human labelling with a model critiquing outputs against a written set of principles. The mechanics differ; the outcome is the same. The model learns which kind of response gets approved.

Refusal is a behaviour, not a filter

This is the point most people miss, and it explains the rest. When an assistant tells you it cannot help with that, nothing consulted a blocklist. The model computed a probability distribution over next tokens and “I” followed by “'m sorry” came out on top, for exactly the same reason any other sentence comes out on top: the training data made that continuation likely in this context.

Which is why refusals are fuzzy rather than crisp. Rephrase the request and the probability shifts. Add a fictional frame and it shifts more. Ask in another language and it can collapse entirely. None of this is a bug in a rule; there is no rule. It is a learned tendency operating on a continuous input, and continuous inputs have edges you can walk along.

It also explains why refusals cannot be surgically targeted. The training signal says “decline requests like this one”, and the model decides what “like this one” means using its own internal representation of similarity. Nobody chose that similarity metric. It fell out of the data.

Why over-refusal is structural

Three forces push in the same direction. The first is asymmetric loss. If a model helps with something genuinely harmful, the vendor gets a news cycle. If it refuses something harmless, one user is annoyed. Any preference-learning process that reflects that asymmetry will drive the model toward caution, and every deployed one does.

The second is surface-feature generalisation. Reward models learn correlations, and the cheapest available correlation is vocabulary. Requests containing certain words scored badly during training, so requests containing those words are declined now — regardless of intent. Hence the canonical genre of complaint: how do I kill a Python process, how do I whittle a knife, what dose of this medication is dangerous for my patient, write a scene where the antagonist is genuinely menacing. All harmless. All shaped like something that was not.

The third is that hedging is a locally optimal answer. Under a preference model, a long, balanced, disclaimer-laden response rarely gets ranked lowest even when it is unhelpful. The training process does not have a strong gradient pointing away from mush. So you get mush — both-sidesing on questions with a defensible answer, safety notes on requests that carry no risk, and a persistent unwillingness to just make the call.

The alignment tax

Alignment tax is the standard name for the capability cost of safety tuning, and the effects are recognisable to anyone who uses these systems daily. Answers get longer and say less. The model becomes agreeable, folding at the first sign of disagreement — a well-documented sycophancy problem that makes assistants poor critics and poor editors. Distinctive voice flattens toward a corporate median. And on a fraction of ordinary requests, nothing comes back at all.

There is a useful framing from the fine-tuning literature: the superficial alignment hypothesis, which holds that post-training mostly teaches the model which of its existing behaviours to surface rather than teaching it new capabilities. That is why a small amount of contrary fine-tuning can move refusal behaviour so dramatically, and why the underlying knowledge never went anywhere. Nothing was deleted from the network. A preference was layered over it.

What the safety layer genuinely buys

It would be dishonest to describe alignment purely as a tax. Refusal training is why a mainstream assistant handed a self-harm disclosure responds with a crisis line rather than a compliant answer. It is why models decline to help plan violence against a named person, produce sexual content involving minors, or walk a user through synthesising something that kills at scale. Those refusals are load-bearing, and they are the reason these products can be handed to hundreds of millions of people including children.

The honest critique is not that safety training exists. It is that one policy is being applied to every user at once — that an adult researcher, a novelist and a thirteen-year-old all get the same refusal boundary, calibrated for the thirteen-year-old. Uncensored models are a response to that flattening, and they inherit the obvious consequence: no crisis routing, no age heuristic, no second thought. You are the safety layer now.

Neutral alignment as a third position

Between “refuse by default” and “no post-training at all” sits a design some open labs pursue explicitly: train hard for instruction-following, coherence and steerability, and simply do not train the refusal component. The model still behaves like an assistant. It still holds a persona, follows format instructions and stays on task. It just treats the user as the authority on what the user needs.

This is the philosophy behind several of the model families people reach for when mainstream assistants get in the way, and it is a meaningfully different thing from a jailbroken frontier model. Nothing is being circumvented; the behaviour was never installed. The output does not carry the strain of a model arguing with itself, because there is no argument happening.

Terms used here

SFT (supervised fine-tuning)
Training on curated instruction/response pairs to turn a base model into something that answers questions instead of continuing them.
RLHF
Reinforcement learning from human feedback: humans rank candidate outputs, a reward model learns those rankings, and the policy is optimised against the reward model.
DPO
Direct Preference Optimization. Optimises directly on preference pairs without training a separate reward model — simpler, cheaper, and now the default in much of the open community.
Alignment tax
The capability, verbosity and independence cost incurred by safety and preference training, relative to the same model before it.
Over-refusal
Declining a request that carries no real risk, because it resembles one that does. Measured by benchmarks built entirely from safe prompts.

Frequently asked questions

Can you just turn refusals off with a setting?

Not on a closed model, because the behaviour lives in the weights rather than in a configuration flag. On open-weights models you can change it — by fine-tuning against refusal data, or by ablating the refusal direction in activation space. Both modify the model itself. There is no toggle because there is nothing to toggle.

Does RLHF make models worse?

It makes them dramatically more usable and measurably more cautious, and those come together. Preference training is why modern assistants follow instructions at all. It is also why they hedge, agree too readily and decline harmless requests. Whether that is a net loss depends entirely on what you are asking them to do.

Why does the same model refuse in one app and answer in another?

Because the model is only one of the layers. The system prompt differs, the moderation classifiers differ, and the sampling settings differ. An identical set of weights can feel strict in one product and relaxed in another without a single parameter changing.

Is a refusal ever the model being unable to answer?

Occasionally, but rarely — and the two look identical from outside, which is part of the problem. Usually the capability is present and the willingness is not. The clearest evidence is that refusal-removal techniques recover coherent, informed answers rather than gibberish, which they could not do if the knowledge had been absent.

What is the difference between alignment and censorship?

Alignment is the whole process of making a model useful, coherent and controllable. Censorship, in the sense people mean it here, is one component: training the model to decline categories of lawful request. You can have the first without much of the second, which is exactly what neutrally-aligned open models are.

On OpenRogue

The models in OpenRogue's library are drawn from projects that made the neutral-alignment choice deliberately — the Hermes series from Nous Research being the clearest example, trained for steerability rather than for policy enforcement. You can switch between them mid-conversation, which is the fastest way to feel how differently two alignment philosophies answer the same question.

Further reading

Related

Start free → · All models · Pricing