The best uncensored AI you can run locally

Your hardware, your weights, your rules — and an honest bill of materials.

Running an uncensored model on your own machine is the strongest answer to the whole category. Nothing leaves the building. No allowance, no account, no terms of service that can change on a Tuesday. The weights on your drive keep working regardless of what any company decides later, which is a form of ownership that hosted AI structurally cannot offer.

The catch is that 'local' is a hardware budget wearing a software costume. The right model is almost entirely determined by how much VRAM you have, so this page is ranked by hardware class rather than by capability — the best local model is the best one that fits.

Worth stating plainly before you read further: this site sells a hosted product, so we have an obvious interest in you choosing the other thing. The page is written to be useful anyway, including the section at the end about where local genuinely wins.

How these were judged

Does it fit

The first and largest filter. A model that does not fit in VRAM either runs at unusable speed off system memory or does not run at all, and no amount of quality makes up for that.

Uncensored at the fine-tune level

Base models are usually lightly aligned and refuse more than people expect. The entries here are community fine-tunes trained without the refusal behaviour, which is a different thing from a base model you have written a clever system prompt for.

Quality per gigabyte

Within a hardware class, which weights actually earn the VRAM. A well-chosen 12B tune beats a badly quantised 70B on a card that cannot really hold it.

Runner ergonomics

Whether you can be chatting in ten minutes or whether the model needs a weekend of configuration. This varies enormously and is rarely mentioned in benchmark tables.

What you give up

Every entry states it, because the honest comparison for local is not against other local models — it is against a hosted service you could be using instead.

This page is about running weights on hardware you own. If you want the same models without the hardware, they are all available hosted, which is what the overall ranking covers. If your interest in local is purely that it costs nothing per message, the free-tier page compares that against the hosted free options directly.

The uncensored models you can run on your own hardware, ranked

The best uncensored AI you can run locally — compared
Rank and nameVRAM classSizeContextEasiest runner
1. Dolphin Venice24GB class, quantised24B32KOllama or LM Studio
2. Mistral Nemo tunes12GB class, quantised12B32KOllama or LM Studio
3. Lunaris8GB class or CPU8B8KOllama or llama.cpp
4. Llama 70B tunesMulti-GPU or heavy quantisation70B131Kllama.cpp or KoboldCpp
5. Magnum v4Multi-GPU or heavy quantisation72B32Kllama.cpp or KoboldCpp
6. WizardLM-2Workstation or server class141B MoE65Kllama.cpp
7. Hermes 3 405BHundreds of GB405B131KNot realistically local

1. Dolphin Mistral 24B Venice Edition

The best capability that fits comfortably on a single high-end consumer card.

Cognitive Computations built the Dolphin line to strip guardrails at training time, and the Venice Edition on Mistral Small 24B is the version that hits the sweet spot for local use. At 24B, quantised, it sits within reach of a single high-end consumer GPU while still being capable enough to be somebody's only model.

It is the local recommendation for most people with a serious card, because it is the largest size where the experience stays pleasant — you are not waiting on tokens, you are not fighting memory pressure, and the model is blunt enough that you stop rephrasing questions to get past anything.

Best at:

Worse at:

Pick this if: Anyone with a high-end consumer GPU who wants one local model and no fuss. · Full Dolphin Venice page

2. Rocinante 12B and UnslopNemo 12B

The mainstream-GPU sweet spot, and the two tunes worth having.

Both of TheDrummer's tunes share a Mistral Nemo 12B base, and 12B quantised is the size that fits a mainstream gaming card without drama. They are ranked together because the choice between them is stylistic rather than technical: Rocinante is the adventurous storyteller, UnslopNemo is the one trained against recycled AI phrasing.

This is the class where local stops being an experiment and starts being usable. Responses arrive fast, the model writes well above its parameter count, and the whole setup fits on hardware a lot of people already own for other reasons.

Best at:

Worse at:

Pick this if: The largest group of people — one mainstream GPU, no interest in a second one. · Full Mistral Nemo tunes page

3. Lunaris 8B (L3)

Runs on almost anything, including hardware that has no business running an LLM.

Sao10K's Llama 3 8B merge is the entry that makes local possible without buying anything. Quantised, an 8B model runs on a modest laptop GPU and will run — slowly but genuinely — on CPU alone, which means the barrier to trying local uncensored AI is an afternoon rather than a purchase.

It is small and it does not pretend otherwise. Blunt short answers, casual scenes, quick questions. As a first local model it is close to ideal, because it teaches you the whole workflow before you have committed any money to it.

Best at:

Worse at:

Pick this if: A first local model, or a laptop that is not going to run anything larger. · Full Lunaris page

4. Euryale 70B and Hermes 3 70B

The ceiling for a serious home rig, and a real step up in quality.

The Llama 70B class is where local output stops feeling like a compromise. Sao10K's Euryale and Nous Research's Hermes 3 70B are the two tunes worth the trouble — one for expressive writing that stays coherent, one for a steerable general assistant with no refusal reflex — and both carry 131K of context if you have the memory to allocate it.

The trouble is real. This class wants multiple GPUs, or aggressive quantisation with the quality loss that implies, or CPU offloading that turns a reply into a coffee break. It is the point where the honest question becomes whether you want to own an inference rig or use one.

Best at:

Worse at:

Pick this if: People who already have the rig, or want one, and value quality over convenience. · Full Llama 70B tunes page

5. Magnum v4 72B

The same hardware class as the 70B tunes, bought for the prose.

Anthracite's Qwen 2.5 72B tune sits in the same hardware bracket as the Llama 70B models and is the reason a lot of people build a rig at all. If you run local specifically to write fiction, this is the target — the prose quality is the best in the open lineup and it is not close.

It ranks below the Llama tunes here only on practicality. Its 32K context is a quarter of theirs, so the memory you free up by running a shorter window is partly the point, and it is a specialist rather than a model you would use for everything.

Best at:

Worse at:

Pick this if: A local rig built primarily for writing fiction. · Full Magnum v4 page

6. WizardLM-2 8x22B

Mixture-of-experts: small-model speed, very large-model memory.

Microsoft published this Mixtral-based mixture-of-experts and withdrew it within days; the weights had already spread and the community has served them ever since. Because it shipped before its final compliance review, it is meaningfully less filtered than anything else with a big-tech pedigree — an accident of process that turned into its main appeal.

Mixture-of-experts routing is widely misunderstood as a memory saving. It is not: all 141B parameters must be resident, and only a fraction are active per token. What you get is the speed of a much smaller model and the memory bill of a very large one, which suits workstations and suits almost nothing else.

Best at:

Worse at:

Pick this if: Workstation owners who want the least filtered large model available locally. · Full WizardLM-2 page

7. Hermes 3 405B

The honest 'you can't' entry — and the clearest argument for hosting.

Nous Research's full fine-tune of Llama 3.1 405B is included precisely because it does not belong on a local list. Running it takes hundreds of gigabytes of VRAM — datacenter hardware, not a home lab — and the heroic quantisation schemes that squeeze it onto less give up much of what made it worth running.

It is here as the boundary of the category. Every other entry on this page is a real choice between local and hosted. This one is the single point where hosted is not a convenience but the only option, and any local guide that omits it is drawing the map with the edges cut off.

Best at:

Worse at:

Pick this if: Nothing, locally. It is here to mark where local ends. · Full Hermes 3 405B page

The runners

The model is half the decision; the software you run it with determines whether the afternoon is pleasant. All of these are free, and none of them is the wrong answer.

Ollama

The fastest path from nothing to chatting. A command-line install, a model library you pull from by name, and sensible defaults. If you have never run a model locally, start here.

LM Studio

A desktop application with a graphical model browser and chat interface. The same job as Ollama with less terminal, which for many people is the whole difference.

llama.cpp and GGUF

The substrate most of the others are built on, and the direct option when you need precise control over quantisation, context and offload split. More work, more control.

KoboldCpp with SillyTavern

The roleplay-oriented stack: a backend that serves the model and a front-end with character cards, personas and fine-grained sampler control. It is the setup most local roleplay guides assume.

Open WebUI

A browser chat interface that sits on top of a local backend, giving you something close to a hosted product's experience while everything stays on your machine.

How to actually choose

Match the model to the card and stop optimising. A 12GB card means a Mistral Nemo 12B tune; a 24GB card means Dolphin Venice; anything in the 70B class means you are building a rig and should know that going in. The most common local mistake is running a badly quantised large model when a well-chosen smaller one would have been better at everything.

Local genuinely wins on three things and it is worth saying so on a site that sells the alternative: your text never leaves your machine, the marginal cost is electricity, and nobody can change the terms. Hosted wins on the ceiling — the 405B class is not coming to a home lab — plus zero setup, phone access, synced history, and not spending an evening reading about quantisation formats. Both of those lists are true, and which one matters more is a question about you rather than about the technology.

Where OpenRogue falls short

Hosted means your text leaves your machine

This is the structural advantage local has and no hosted service can neutralise it. If that is your requirement, local is the answer and the rest of the comparison is noise.

No offline mode

OpenRogue needs a connection. Local models do not, and on a plane or behind a firewall that difference is total.

Hosted is a recurring cost

Local costs a lot once; hosted costs a little forever. Over a long enough horizon and with hardware you already own, local is cheaper and pretending otherwise would be silly.

Less control over sampling

A local stack exposes temperature, repetition penalties, sampler order and prompt templates. A hosted chat product exposes far less of that surface, and power users notice.

Frequently asked questions

What is the best uncensored AI to run locally?

It depends almost entirely on your VRAM. On a 12GB card, Rocinante 12B or UnslopNemo 12B; on a 24GB card, Dolphin Mistral 24B Venice Edition; on a multi-GPU setup, Euryale 70B or Hermes 3 70B. Choosing the largest model that genuinely fits beats forcing a bigger one into too little memory.

How much VRAM do I need for uncensored AI?

An 8B model quantised runs on an 8GB card and will even run on CPU. A 12B needs roughly a mainstream 12GB card, a 24B wants a 24GB card, and the 70B class needs multiple GPUs or aggressive quantisation. Context length costs memory on top of the weights, so budget for both.

Is running AI locally actually private?

Yes, in the strongest sense available. The model runs on your hardware, so no prompt or response is transmitted anywhere and no provider policy applies to it. That is the one advantage no hosted service can match, including this one.

Can I run a 405B model at home?

Realistically, no. Hermes 3 405B needs hundreds of gigabytes of VRAM, which is datacenter hardware. Quantisation schemes that force it onto less give up much of what makes it worth running, so this is the one place where a hosted service is the only practical route.

What is the easiest way to start running models locally?

Install Ollama or LM Studio, pull an 8B or 12B uncensored tune, and chat. Both handle quantisation and defaults for you, and starting small means you learn the workflow before you have spent anything. Move to llama.cpp directly when you need control the wrappers do not expose.

Related

Start free → · All models · Pricing