A language model does not produce text. It produces a probability distribution over the next token, and does it again for every token after that. Everything you experience as the model's personality — precision, flair, repetition, the occasional confident nonsense — is decided by the small pile of settings that choose one token from that distribution.
Controls · 8 min read
Generation is a loop. The model reads the context and outputs one raw score — a logit — for every token in its vocabulary. Those scores pass through the sampler stack, which modifies and filters them, and then through a softmax that turns them into probabilities. One token is drawn, appended to the context, and the loop runs again for the next.
Every setting in the list below operates in that gap between the logits and the draw. None of them change the model, and none of them add information. They change which of the model's own candidate continuations you are likely to see — which is a much bigger lever on perceived quality than most people assume.
Temperature divides the logits by a constant before the softmax. Divide by a number below one and the gaps between scores widen, so the leading candidate takes even more of the probability mass and output becomes more deterministic, more conventional, more repetitive. Divide by a number above one and the gaps compress, so unlikely tokens become genuinely reachable and output becomes more varied, more surprising, and eventually incoherent.
At temperature zero, sampling collapses to always taking the highest-scoring token — greedy decoding. In practice that is still not perfectly reproducible on GPU hardware, because floating-point non-determinism in batched inference can flip near-ties, so treat zero as “as deterministic as this system offers” rather than as a guarantee.
The failure modes sit at both ends and look completely different. Too low: the model repeats phrasings, falls into structural tics, reaches for the same three metaphors and produces prose with no surprises in it. Too high: sentences start well and lose their thread, names drift, facts wobble, and the model invents vocabulary. If output feels stale, raise it; if output feels unreliable, lower it before you blame the model.
Truncation samplers exist to make higher temperatures survivable. They cut the tail off the distribution before sampling, so the genuinely absurd candidates cannot be drawn no matter how much the temperature flattened things.
Top-k keeps the k highest-scoring tokens and discards the rest. Simple and blunt: the same k is applied whether the model is completely certain or genuinely torn, which is the wrong behaviour in both cases.
Top-p, also called nucleus sampling, keeps the smallest set of tokens whose probabilities sum to at least p. It is adaptive, which is why it became the default — when the model is confident the set is tiny, when it is uncertain the set widens. It still misbehaves at the extremes: a very flat distribution can pull a large number of mediocre candidates into the nucleus.
Min-p takes a different approach: it sets the cutoff relative to the top token's probability, keeping only tokens above some fraction of the leader. That scales naturally with confidence and tends to hold up better at high temperature, which is why it has become popular in the local writing and roleplay community for creative work. Running all three at once is usually redundant; pick one truncation sampler and tune it.
Repetition penalty reduces the score of tokens that already appeared, making them less likely to recur. Frequency and presence penalties are the same idea split in two: one scales with how often a token appeared, the other applies a flat penalty for having appeared at all. DRY — don't repeat yourself — is a more targeted variant that penalises the continuation of repeated multi-token sequences rather than individual tokens, which is much better suited to fiction where names and terms must recur freely but phrasings should not.
The trap is that these penalties do not understand language. A blunt repetition penalty applied hard will suppress a character's name, common function words, and the correct technical term, forcing the model toward strained synonyms. Text written under an aggressive penalty has a recognisable quality: nobody is called the same thing twice.
Sustained looping is usually a symptom rather than a setting problem. Temperature far too low, a context saturated with near-identical turns, or a model pushed past its trained context length will all produce loops that no penalty properly fixes. Diagnose those first.
The samplers form a pipeline, and the pipeline order is not standardised. Some stacks apply temperature before truncation, some after, and some expose the ordering as a setting. The difference is real: truncating first and then flattening the survivors behaves quite differently from flattening everything and then truncating.
This is the main reason a settings recipe copied from a forum does not reproduce. Identical numbers in two applications can be two different pipelines, on top of different default values for the parameters nobody mentioned. Tune in the tool you are actually using, change one thing at a time, and treat any recipe as a starting point.
These are conventions rather than measurements, and they are meant as a first guess to adjust from. For factual work, summarisation, extraction and code, use a low temperature — the region around 0.2 to 0.4 — with a fairly tight nucleus. You want the model's best guess, not its imagination.
For general conversation, something near the middle, around 0.7 to 0.8, with top-p around 0.9. This is roughly where most chat defaults sit for good reason.
For fiction, roleplay and anything where voice matters, higher — 0.9 to 1.1 is common — paired with a truncation sampler to keep the tail from doing damage, and a repetition control that penalises phrases rather than tokens. If the prose is flat, the temperature is usually the reason. If names and continuity are slipping, it is usually too high.
It cannot add knowledge. If the model does not know something, no temperature setting will retrieve it — a high temperature simply makes the invention more varied, and a low one makes the same wrong answer arrive with more conviction. Confident phrasing is a property of the sampler, not of correctness.
It cannot reliably remove a refusal either. If declining is the dominant continuation, raising the temperature only re-rolls the same distribution; you may occasionally sample a compliant opening, and the refusal-trained model will frequently steer back mid-paragraph. Refusal is a property of the weights, which is the whole reason the uncensored-model ecosystem exists rather than a folklore of clever settings.
Somewhere around 0.9 to 1.1 is the usual starting point, paired with a truncation sampler such as top-p near 0.9 or a min-p cutoff. Adjust from there by symptom: flat, predictable prose means raise it; drifting names, wobbling facts or sentences that lose their thread mean lower it.
Top-k always keeps a fixed number of candidates regardless of how confident the model is. Top-p keeps however many candidates are needed to reach a cumulative probability threshold, so the set shrinks when the model is certain and widens when it is not. Top-p is adaptive, which is why it is the more common default.
Nearly, but not guaranteed. Greedy decoding always takes the top-scoring token, so the algorithm is deterministic — but floating-point non-determinism in batched GPU inference can flip near-ties between runs. Treat it as maximally consistent rather than truly reproducible.
Most often temperature set too low, which makes the likeliest continuation overwhelmingly likely and collapses the output into loops. Other common causes are a context filled with near-identical turns, or exceeding the length the model was trained on. Try raising temperature and trimming the context before reaching for a heavy repetition penalty.
Not reliably. Refusal is what the model's distribution favours in that context, and sampling only chooses among the candidates it already produced. You may occasionally draw a compliant opening at high temperature, and a refusal-trained model will often correct course mid-response. Removing refusal requires changing the model, not the sampler.
Sampling explains a lot of what people attribute to a model's personality, which is worth knowing before switching models to fix a problem the settings caused. That said, some things genuinely are the model — UnslopNemo's fresher word choice comes from its training data, not from a temperature setting, and no sampler will make a refusal-trained model comfortable with your scene.