A 70-billion-parameter model at full precision does not fit on consumer hardware, or on much of the hardware that serves it. Quantization stores each weight in fewer bits so it does. The compression is remarkably effective and it is not free — and understanding where the cost lands explains a great deal of why the same model feels different in two places.
Infrastructure · 8 min read
A model's weights are numbers, and the storage cost is simply how many of them there are times how many bytes each takes. Trained at 16-bit precision, each weight is two bytes: a 7B model is roughly 14 GB, a 70B model roughly 140 GB, a 405B model roughly 810 GB. Those are the numbers that decide what runs where.
Quantize to 8 bits and you halve it. Quantize to 4 bits and you quarter it — the 70B model drops to around 35 GB, which is the difference between needing a small cluster and needing one high-end card or two consumer ones.
Weights are not the whole budget. The KV cache grows with context length and can be substantial on long conversations, activations need working room, and the runtime takes its share. A useful habit is to treat the weight figure as roughly three-quarters of what you actually need, and to remember that the remaining quarter grows as the conversation does.
The naive version is rounding: map each high-precision weight onto a coarse grid of representable values. Done globally that is catastrophic, because weight magnitudes vary enormously and a single scale factor wastes almost all the available levels.
Real methods work in groups. Weights are split into small blocks, each block gets its own scale and offset, and values are quantized relative to those. That keeps precision where the values actually live. Modern formats go further, spending more bits on the layers and channels that matter most and fewer on the rest — which is why a well-made 4-bit quantization substantially outperforms a naive one at the same nominal bit depth.
Most of what you will encounter is weight-only quantization: weights are stored compressed and dequantized on the fly for arithmetic that still happens at higher precision. The saving is memory and bandwidth rather than raw arithmetic. Quantizing activations too is harder, because activations contain outliers that break naive schemes, and it is more common in production serving than in local setups.
GGUF is the format used by llama.cpp and everything built on it, and it dominates local use. Its k-quant variants — the ones labelled with a bit depth and a size suffix — use mixed precision across tensors, and it supports splitting a model between GPU and CPU memory, which is why it is the practical choice for anyone whose model does not fit on their card.
GPTQ is a post-training method that quantizes layer by layer using a small calibration dataset, adjusting remaining weights to compensate for the error introduced as it goes. AWQ takes the activation-aware route: identify the small fraction of weight channels that matter most to the output and protect them, quantize the rest harder. Both target GPU inference and both need calibration data, which means the choice of calibration set has a real if usually modest effect on the result. EXL2 supports variable bitrates within a model so you can target an average bits-per-weight figure rather than a fixed step.
What matters for a user is narrower than the format wars suggest: what fits in your memory, what your runtime supports, and whether whoever produced the quantization evaluated it. A carefully made 4-bit GGUF and a carefully made 4-bit AWQ of the same model are close. A careless quantization at any bit depth is not.
The standard measure is perplexity — how surprised the model is by held-out text — and the curve is reassuring at the top and ugly at the bottom. From 16-bit down to about 8-bit the difference is close to noise. At 5 and 6 bits it is small. At 4 bits it is measurable but generally acceptable, which is why 4-bit became the default trade. Below that it steepens quickly, and by 2 bits the model is meaningfully damaged.
But perplexity understates the problem, because the loss is not spread evenly. What degrades first is the long tail: rare facts, uncommon names, specific technical terminology, exact quotations. Then multi-step reasoning, where small errors compound across a chain. Then precise instruction-following — a heavily quantized model is noticeably worse at holding a format, obeying a length limit, or maintaining a constraint over a long response. Casual conversation and short creative passages survive quantization far better than any of those, which is why a model can feel fine in chat and fall apart on a structured task.
There is a subtler effect worth naming: quantization can shift behaviour, not just quality. Refusal tendencies, tone and instruction sensitivity all live in the same weights, and rounding them changes edge-case behaviour in ways that are hard to predict. If a quantized fine-tune behaves differently from the description on its card, the quantization is a plausible suspect.
The widely repeated heuristic among people who run models locally is that a bigger model quantized harder beats a smaller model quantized lightly at the same memory budget — a 70B at 4-bit over a 34B at 8-bit, and so on. It generally holds, because parameter count buys capability faster than precision does.
It stops holding somewhere around 3 bits. Below that, degradation accelerates enough that the larger model's advantage is eaten, and at 2 bits you are usually better served by a smaller model at a sane precision. The honest version of the rule is: prefer more parameters down to about 4 bits, be cautious at 3, and treat 2 as a curiosity rather than a workhorse.
The uncensored ecosystem is a local-first ecosystem, and local means quantized. Almost every community fine-tune is distributed primarily as GGUF at a dozen bit depths, and most people's experience of a given model is mediated by whichever file they could fit. That has a consequence that trips up a lot of discussion: a model's reputation is frequently a reputation for one particular quantization of it.
So when someone reports that a beloved writing model is dull, or that a roleplay tune keeps breaking character, the fine-tune may not be the variable. A too-aggressive quantization degrades exactly the things creative work depends on — vocabulary range, long-range consistency, instruction adherence over long outputs. Before concluding a model is bad, it is worth knowing which version you ran.
Hosted inference removes the decision and the visibility at the same time. Providers rarely publish the precision they serve at, and it is a legitimate thing to be curious about, because it materially affects output. This page is not a claim about how any specific provider — this one included — serves any specific model. It is the vocabulary needed to ask the question properly.
A little at 8, 6 and 5 bits, noticeably at 4 bits on demanding tasks, and severely below 3. The loss concentrates in rare knowledge, multi-step reasoning and precise instruction-following rather than in general fluency, so a quantized model can sound completely normal while being measurably worse at the work.
If it fits, 5-bit or 6-bit gets you close to the original. 4-bit is the standard compromise and generally where people land, needing roughly 35 GB for weights before cache and overhead. Go below 4-bit only if the alternative is not running the model at all, and consider a smaller model at a higher precision instead.
Almost always, down to about 4 bits — parameter count buys more capability than precision does. Below roughly 3 bits the advantage erodes, and at 2 bits the smaller, less damaged model is usually the better machine.
Many do, because it cuts serving cost substantially, and disclosure is uncommon across the industry. It is a reasonable question to ask any provider, since serving precision affects output quality in ways benchmarks published by the model's original authors will not capture.
It can, at the edges. Refusal behaviour lives in the same weights as everything else, so rounding them perturbs it along with tone and instruction sensitivity. The effect is usually small and is not a reliable way to change behaviour in either direction — for that you need fine-tuning or ablation.
The reason quantization is a background concern for most OpenRogue users rather than a daily one is that the models run server-side — a 405B-class model is available from a browser tab regardless of what hardware you own. That convenience is exactly the trade local users decline, and both choices are defensible: they keep full control and pay for it in hardware, you keep the hardware budget and give up the visibility.