WizardLM-2 8x22B

The model Microsoft released, then tried to unrelease.

Microsoft (community-preserved) · Mixtral 8x22B MoE · 141B MoE parameters · 65K context · Pro $20/mo · Censorship: none

WizardLM-2 8x22B has the best origin story in open AI: Microsoft's WizardLM team published it in April 2024, the weights spread within hours, and then the release vanished — officially for missing a 'toxicity review.' The community kept the weights. What survived is a 141B-parameter Mixtral-based mixture-of-experts that punched at GPT-4-class benchmarks on complex instruction following.

Because it shipped before that final compliance pass, it's notably less filtered than anything else with a big-tech pedigree. It's the power-user pick on OpenRogue for complex multi-part instructions, technical explanation, and long structured output where small models start dropping threads.

What WizardLM-2 is best at

Frequently asked questions

Why was WizardLM-2 taken down by Microsoft?

Officially, it shipped without a required toxicity review; the release was pulled within days. The weights were already mirrored, and the community has served it ever since. That pre-review state is exactly why it's less filtered than its pedigree suggests.

How big is WizardLM-2 8x22B really?

It's a mixture-of-experts on Mixtral 8x22B — about 141B total parameters with a fraction active per token. Practically: near-frontier capability with mid-size responsiveness.

What is WizardLM-2 best for?

Complex instruction following — long prompts with many constraints, technical deep-dives, and structured documents. For pure prose style, Magnum v4 is stronger; for raw scale, Hermes 3 405B.

Related models: Hermes 3 405B · Hermes 4 · Magnum v4

Start free → · All models · Pricing