Open weights vs open source AI

Three phrases get used interchangeably and mean genuinely different things. One describes what you can download, one describes what you are permitted to do with it, and one describes how much of the process you can inspect. Getting the distinction right matters practically, because only one of these categories can be modified — and modification is the entire basis of uncensored AI.

Fundamentals · 7 min read

In short

Three things people mean by open

Open access is the weakest form: you can use the model through an API. You cannot download it, inspect it, run it offline or change it. Every frontier commercial assistant is open access at best, and calling it open at all is a stretch.

Open weights is the meaningful middle. The parameter file is published, usually on a model hub, and you can download it, run it on your own hardware, fine-tune it, merge it, quantise it and redistribute the result — subject to a licence. This is the category that makes the entire open ecosystem exist. Almost everything people call “open source AI” is actually this.

Open source, used with precision, is stricter. The Open Source Initiative's definition for AI systems asks for the parameters, the code used to train and run the system, and sufficiently detailed information about the training data that a competent party could recreate a substantially equivalent system — all under terms that permit use, study, modification and sharing without restriction on field of endeavour. Very few widely used models meet that bar. The distinction is not pedantry: it is the difference between being allowed to use something and being able to understand it.

What the licences actually restrict

Some model families ship under genuinely permissive terms — Apache 2.0 or MIT — which impose little beyond attribution and warranty disclaimers. Others ship under bespoke community licences that are open-ish in a specific, bounded way, and the details vary enough that reading the actual text is the only reliable approach.

Recurring clauses to look for: an acceptable use policy listing prohibited applications; a scale threshold above which very large deployers must negotiate separate terms; naming or attribution requirements on derivative models; restrictions on using outputs to train competing models; and occasionally field-of-use carve-outs. None of these make a licence bad. They do make “open source” the wrong word for it.

The nuance that trips people up: an acceptable use policy is a contract between you and the model's publisher. It is not enforced by the weights, does not disappear when you fine-tune, and applies to you as the deployer whatever behaviour your derivative exhibits. A fine-tune that removes refusals removes a behaviour, not an obligation. Those are separate layers and they stay separate.

Why this decides who can be uncensored

Every technique for removing refusal behaviour needs the weights. Fine-tuning computes gradients through them. Abliteration reads internal activations and rewrites projection matrices. Merging combines them. All of it requires possession of the parameters, which open access categorically does not provide.

That is why the uncensored ecosystem is built entirely on open-weights releases, and why it clusters around a handful of base families — the ones large enough to be worth the effort and permissively licensed enough to redistribute. When a lab ships a strong open-weights model, a wave of fine-tunes follows within weeks. When a lab ships an API, nothing follows, because there is nothing to build on.

It also explains the shape of the market. Closed frontier models lead on capability and cannot be modified. Open models trail on the hardest reasoning and can be reshaped completely. Choosing between them is choosing which of those two properties you need today, and the answer differs by task.

Open data is the piece nobody ships

Weights without data are a compiled binary without source. You can run it, benchmark it and modify it by trial and error, but you cannot answer basic questions: what is in it, what is missing, what was filtered, whose work is inside, which biases were baked in at pretraining and therefore cannot be tuned away.

A small number of projects do release corpora and full training pipelines — the OLMo and Pythia efforts are the usual examples — and they are disproportionately valuable to researchers for exactly this reason. Most flagship open-weights models do not, generally for a mix of competitive and legal reasons. Which means the honest description of today's ecosystem is: open weights, closed data, and a lot of vocabulary papering over the gap.

For an uncensored-model user this has one concrete consequence. When a model is oddly ignorant about a subject, or oddly opinionated, or refuses something even after refusal training was removed, the cause may be pretraining data curation — and no amount of fine-tuning fully reverses what was never learned.

Reading a model card without being fooled

Four questions get you most of the way. Which base model is this, and what is its licence? What was the fine-tuning method — SFT, DPO, merge, abliteration — and on roughly what kind of data? What context length was it actually trained at, as opposed to what the config file will accept? And has anyone other than the author evaluated it?

Two warning signs. A card that lists benchmark wins with no methodology is decoration. And a card that describes a model as “fully unrestricted” without describing what was done to it is usually either an abliteration someone did not evaluate, or a merge of other people's work with a new name attached. Neither is disqualifying. Both mean the evaluation is your job.

Terms used here

Open weights
The trained parameters are published and downloadable. Says nothing about licence terms, training data disclosure, or the code used to produce them.
Open access
The model is usable only through an API. Not downloadable, not inspectable, not modifiable. Every frontier commercial assistant is in this category.
Acceptable use policy
A list of prohibited applications attached to a model licence. Contractual, binding on the deployer, and unaffected by fine-tuning the model's behaviour.
Community licence
A bespoke licence that permits most use and redistribution while adding conditions — typically an AUP, a scale threshold, and derivative-naming requirements.
Model merge
Combining the weights of two or more models arithmetically to blend their behaviours, without any additional training. Common in the open roleplay and writing community.

Frequently asked questions

Is Llama open source?

It is open weights under a community licence, which is not the same thing. The weights are downloadable and broadly usable, but the licence adds conditions — an acceptable use policy, naming requirements for derivatives, and terms that apply to very large-scale deployers — and the training data is not published. Strictly speaking that fails the open-source definition while remaining genuinely useful.

Can I use an open-weights model commercially?

Usually, but it depends on the specific licence and you have to read it. Apache 2.0 and MIT models are straightforward. Community licences typically permit commercial use with conditions attached. A few research-only releases prohibit it outright. Never infer the licence from the fact that the download worked.

Does fine-tuning a model let me ignore its licence?

No. Derivative works inherit the original terms, which is why licences include naming clauses for derivatives. Removing refusal behaviour changes what the model does; it does not change what you agreed to when you downloaded the base.

Why does open data matter if I just want to chat?

Mostly it does not, until the model behaves oddly. Gaps and biases from pretraining cannot be fine-tuned away and are invisible without data disclosure. When a model is strangely ignorant or strangely insistent about a subject, closed data is why nobody can tell you exactly why.

Which is better, an open or a closed model?

For the hardest reasoning, coding and multimodal work, the strongest closed models still lead. For anything requiring modification, self-hosting, guaranteed availability, or freedom from a content policy, open weights win by default — because those things are impossible otherwise.

On OpenRogue

Every model in OpenRogue's library is an open-weights release or a community fine-tune of one, which is not a philosophical stance so much as a structural requirement — a closed model cannot be uncensored, only argued with. Each model page names the base model, the creator and the parameter count so you can go and read the original card yourself.

Related

Start free → · All models · Pricing