Ten open models trained without a refusal layer — plus every frontier model, in the same dropdown.
Most AI products are one model with a marketing department attached. This is a library: ten genuinely uncensored open models, each fine-tuned by a different group with a different obsession — prose quality, reasoning, speed, roleplay, the eradication of AI clichés — alongside the frontier models from the GPT, Claude, Gemini and Llama families on the tiers that include them.
They are not interchangeable, and the interesting part of using OpenRogue is learning which one to reach for. Every model below has its own page with specs, what it is actually good at, sample prompts, and a straight answer about when something else in the lineup would serve you better.
Every open model here was fine-tuned without refusal training. That is a property of the weights, not a prompt trick: there is no safety layer to argue with, so the model does not stop mid-scene to explain what it is not comfortable with, and it does not need the question rephrased three times before it engages.
OpenRogue adds nothing on top of them — no classifier, no injected policy, no topic blocklist. The frontier models on the higher tiers are a different matter and keep whatever policies their vendors trained into them; that is precisely why the uncensored library exists next to them rather than instead of them.
Sizes run from an eight-billion-parameter model that replies faster than you can type up to a four-hundred-billion-parameter fine-tune that is the largest neutrally-aligned model you can chat with online. Context windows vary just as much, which matters more than parameter count once a session gets long.
The shortest useful version, by the job rather than by the benchmark. All of these are one click apart, so treating the first choice as reversible is the right instinct.
Dolphin Mistral 24B Venice Edition. Fast, blunt, free, and the default for a reason — if you are here to find out what uncensored actually feels like, start there.
Magnum v4 72B, trained explicitly toward frontier-grade writing quality, which commits to a scene instead of fading out of it.
Euryale 70B for the balance of expressiveness and intelligence, with a long context window for campaigns that run for weeks.
Hermes 4 70B, which reasons step by step on demand and pairs with Thinking mode, or Hermes 3 405B when depth matters more than latency.
WizardLM-2 8x22B — a mixture-of-experts whose benchmark speciality was following long prompts with many constraints.
UnslopNemo 12B, fine-tuned against the recycled phrasing that makes AI text obvious at a glance.
Lunaris 8B or Rocinante 12B, for rapid exchanges, quick questions and phone sessions where waiting kills the rhythm.
The higher tiers add the frontier models from the GPT, Claude, Gemini, Llama and Qwen families to the same dropdown. They arrive with their vendors' policies intact — nothing here removes those, and any product claiming otherwise is describing a jailbreak.
The reason to have them alongside the uncensored library is that they are genuinely better at some things, and switching costs one click. Reason through the specification on a frontier model, then move the same thread to an uncensored model for the part it would decline. That workflow is the actual argument for a multi-model product.
Context is how much of the conversation the model can still see. Below a certain size, a long roleplay quietly forgets its own plot and a document analysis starts answering from the last few pages only — which reads as the model getting dumber, when it has simply stopped being shown the beginning.
The larger models in this library carry substantially longer context windows than the small fast ones, and each model page lists its own. If a session has started losing the thread, moving it to a longer-context model is usually the fix, and it is the same one click as any other switch.
Nothing here is a commitment. Change the model from the picker and the conversation continues, with the new model reading everything that came before. Draft fast, escalate for the hard passage, drop back down when the scene turns light again — or send one prompt to several models at once and compare the answers side by side before choosing.
Dolphin Mistral 24B Venice Edition. It is the default model, it is free to use, and it is blunt enough that you will know within a handful of messages whether uncensored AI changes your work. From there, move toward Magnum or Euryale for creative work and the Hermes family for reasoning.
They were fine-tuned without refusal behaviour, which is a property of the weights rather than a prompt. That is why the behaviour stays consistent instead of degrading over a long session, and why it cannot be patched out from under you the way a jailbreak can.
Yes, on the tiers that include them, in the same dropdown as the uncensored library. They keep their own vendors' content policies — no platform can remove those — so the practical pattern is to use them for what they are best at and switch to an uncensored model for anything they decline.
No. The thread continues and the new model reads the full history that came before it. That is the whole reason to keep several models behind one interface rather than several subscriptions behind several tabs.
No. Running a four-hundred-billion-parameter model locally takes hundreds of gigabytes of VRAM; here it runs in the cloud and you pick it from a dropdown. The tradeoff is honest and worth stating: local models are private in a way no hosted service can match, and hosted models are large in a way no consumer machine can match.