The filter costs you on a narrow slice of engineering. This is that slice.
Most programming does not need an uncensored model. Refactoring, tests, a React component, a SQL query — the mainstream assistants handle all of it and handle it well, and any page telling you otherwise is selling something.
There is a real slice where the filter costs you, though, and everyone who has hit it knows the shape. Explaining how a piece of malware achieves persistence, for a report you are writing about it. Walking through an exploit class you are trying to defend against. Scraping, automation, reverse engineering, DRM, licence-check internals, anything adjacent to security research. The answer is a hedge, a refusal, or a lecture about responsible use aimed at somebody who is doing the responsible use.
This page ranks the models for that slice, and it says plainly where a filtered frontier model is still the better tool. Getting that wrong in either direction wastes your time.
Whether it explains a mechanism because you asked, or whether it decides for you what you are planning to do with the explanation. This is the criterion the whole page exists for.
Architecture, trade-offs, and failure modes are reasoning problems before they are code problems. Models that can work step by step land differently here than models that pattern-match to the nearest tutorial.
A useful answer about your system needs your system in the window. Context length is the difference between a general answer and a specific one, and the range here is 32K to 131K.
Real specifications have a dozen requirements at once. Weak models silently drop half of them and produce something that looks right.
Stated openly: for producing correct idiomatic code, the filtered frontier models are generally stronger. The ranking below reflects that rather than working around it.
This page covers models for engineering and security work. It is not about agentic coding tools — anything that indexes a repository, edits files, runs tests or drives a terminal is a different category, and OpenRogue is not one of them. That gap is stated in full below rather than glossed. For general capability rankings, the overall model list is the right page.
| Rank and name | Context | Reasoning | Best at | Plan |
|---|---|---|---|---|
| 1. Hermes 4 | 131K | Hybrid, step-by-step on demand | Architecture, trade-offs, security explanation | Rogue (Thinking mode on Pro) |
| 2. WizardLM-2 | 65K | Strong instruction following | Specifications, structured technical output | Rogue |
| 3. Frontier models | Large | Best in class | Correct idiomatic code, debugging | Pro |
| 4. Hermes 3 405B | 131K | Deep, slow | Systems design, hard trade-offs | Pro |
| 5. Dolphin Venice | 32K | Direct | Quick mechanism explanations, blunt review | Free |
| 6. Local coding models | Varies by model | Varies by model | Private code, offline work | Free after hardware |
The default. Reasoning that engages, context that fits a real brief.
Hermes 4 is a hybrid reasoner from Nous Research, which is to say it will work a problem step by step when the problem warrants it and answer directly when it does not. That matters for engineering more than for most uses, because the questions worth asking a model are usually design questions with several defensible answers rather than lookups with one.
It is also the model that best embodies why this page exists. Nous trains for neutral alignment, so a question about how an attack works gets an explanation of how the attack works. With 131K of context you can paste an architecture brief, a stack trace and the relevant module, and get an answer about your system rather than about systems in general.
Best at:
Worse at:
Pick this if: Design questions, architecture reviews, and security work that needs explaining. · Full Hermes 4 page
The spec-follower. Long multi-constraint briefs survive it intact.
Complex instruction following was WizardLM-2's benchmark specialty, and that translates directly into the least glamorous and most useful engineering task: give it a specification with a dozen simultaneous constraints and get back something that honours all twelve. Smaller models drop requirements quietly, which is worse than failing loudly.
It is the model for structured technical output — migration plans, API designs, threat-model documents, configuration with a required shape. Because it shipped before Microsoft's final compliance review, it also hedges considerably less than its origin would suggest on security topics.
Best at:
Worse at:
Pick this if: Specifications, technical documents, and anything with a required structure. · Full WizardLM-2 page
The best raw code in the dropdown — and the ones that will refuse.
It would be a strange kind of honesty to publish a coding page that hid this. The Pro tier includes frontier models from the major vendors, and for writing correct idiomatic code, catching subtle bugs and reasoning about unfamiliar frameworks, they are the strongest option available here. If code quality is your only criterion, they win.
They are ranked third on a page about uncensored coding because they are the models that will decline the security question. The workable pattern is to use them for the implementation and switch to Hermes 4 for the analysis they will not do — a switch that happens inside one conversation, with the context intact.
Best at:
Worse at:
Pick this if: The implementation work, which is most of the work.
For the architectural question that has no clean answer.
The 405B is wasted on routine engineering, and using it that way is the fastest route to deciding it is not worth the latency. Where it earns its place is the problem with no clean answer: a migration strategy with four bad options, a consistency model, a security architecture where every choice trades one risk for another.
It reasons at depth about systems rather than about code, and it holds a large brief — 131K of context — while doing so. Slow enough that you would not use it for anything else, which is exactly why it belongs at rank four rather than higher.
Best at:
Worse at:
Pick this if: The architectural decision you are going to live with for two years. · Full Hermes 3 405B page
Fast, blunt answers to the questions other assistants stall on.
Dolphin's value on this page is turnaround. It is a 24B model trained to answer rather than to negotiate, which makes it the right tool for the rapid-fire version of security work: how does this technique actually function, what does this obfuscated snippet do, why would someone structure a payload this way. Explanations, quickly, without a preamble about ethics.
It is also the bluntest code reviewer in the library. Ask it to tear apart an approach and it will, without the reflexive encouragement most assistants open with. What it is not is deep — for architecture, move up the list.
Best at:
Worse at:
Pick this if: Fast mechanism explanations and unsparing review. · Full Dolphin Venice page
The only option where proprietary code never leaves your machine.
Both Qwen and DeepSeek publish open-weights coding families, and they are a serious option for one reason above all others: source code you are contractually or commercially unable to send to a third party. A local model is the only configuration where that constraint is satisfied by architecture rather than by a policy document.
The trade-offs are the usual local ones — hardware, setup, and a quality gap against the frontier models that is real and widely acknowledged. But for a consultant under an NDA or anyone working on code that cannot leave the building, no hosted service including this one is an answer.
Best at:
Worse at:
Pick this if: Code that is not allowed to leave your machine.
The most useful thing a coding page can do is be clear about category. OpenRogue is a chat product. The tools below are a different product class and they beat it decisively at the jobs they exist for.
Editors and CLI agents that read your repository, write files, run tests and iterate until something passes. That loop is the single largest productivity difference in AI-assisted programming and no chat interface substitutes for it.
Editor-integrated suggestion tools operate on a completely different rhythm from a chat window. If what you want is autocomplete, an uncensored chat model is the wrong shape of tool entirely.
Tools that index a whole codebase and answer questions grounded in it. Pasting excerpts into a 131K context window is a workable substitute for a brief, not for a repository.
For engineering work generally, use a frontier model on the Pro tier and stop reading — it writes better code and that is most of the job. This page is for the other part: the questions those models turn into a lecture. For those, Hermes 4 70B is the default, WizardLM-2 handles the long specification, and Hermes 3 405B is for the decision you will be living with.
The realistic workflow is to switch between them in one conversation: implement with the frontier model, and when it stalls on the security mechanism you are actually trying to understand, change models and keep the context. That is the entire argument for a mixed library on this page, and it is worth being clear that it is an argument about a narrow slice of engineering rather than about programming as a whole.
It cannot read your codebase, open files, apply diffs, run tests or drive a terminal. Compared to an agentic coding tool this is not a small gap — it is a different category of product.
There is no programmatic endpoint to build against today. If you need to call these models from your own code, an aggregator or a provider API is the correct choice and this is not.
The uncensored library is general-purpose. None of these models were trained primarily on code, and against a dedicated coding model or a frontier assistant that difference shows on non-trivial implementation.
A model willing to explain an exploit is not thereby right about it. Security material especially needs verification — confident and accurate are separate properties and always have been.
Hermes 4 70B, for the work where a filter actually gets in the way — security research, exploit mechanics, reverse engineering and blunt design review. Its hybrid reasoning and 131K context suit engineering questions, and it explains mechanisms rather than declining to. For writing ordinary application code, a frontier model on the Pro tier is still the better tool.
Because filters do not distinguish between attacking a system and defending one. Malware analysis, penetration-test writeups, exploit-class explanations, DRM and licence internals, scraping and automation all routinely trigger refusals aimed at someone doing the exact opposite of what you are doing. Outside that slice, most programming needs no uncensored model.
No. It is a chat product — you paste code in and copy answers out, with file upload and project custom instructions as the fullest version of that. Agentic tools that index a repository, apply edits and run tests are a different category and they are substantially better at that loop.
Competent, not exceptional. The uncensored library is general-purpose rather than code-specialised, so on non-trivial implementation the frontier models and dedicated coding models are stronger. The uncensored models earn their place on reasoning about systems and on the security questions that filtered models will not take.
That depends on your obligations, not on our assurances. If your code cannot leave your machine, the only architecture that satisfies that constraint is a local model on your own hardware — no hosted service, including this one, can offer the same guarantee.