The model Microsoft released, then tried to unrelease.
Microsoft (community-preserved) · Mixtral 8x22B MoE · 141B MoE parameters · 65K context · Pro $20/mo · Censorship: none
WizardLM-2 8x22B has the best origin story in open AI: Microsoft's WizardLM team published it in April 2024, the weights spread within hours, and then the release vanished — officially for missing a 'toxicity review.' The community kept the weights. What survived is a 141B-parameter Mixtral-based mixture-of-experts that punched at GPT-4-class benchmarks on complex instruction following.
Because it shipped before that final compliance pass, it's notably less filtered than anything else with a big-tech pedigree. It's the power-user pick on OpenRogue for complex multi-part instructions, technical explanation, and long structured output where small models start dropping threads.
Officially, it shipped without a required toxicity review; the release was pulled within days. The weights were already mirrored, and the community has served it ever since. That pre-review state is exactly why it's less filtered than its pedigree suggests.
It's a mixture-of-experts on Mixtral 8x22B — about 141B total parameters with a fraction active per token. Practically: near-frontier capability with mid-size responsiveness.
Complex instruction following — long prompts with many constraints, technical deep-dives, and structured documents. For pure prose style, Magnum v4 is stronger; for raw scale, Hermes 3 405B.
Related models: Hermes 3 405B · Hermes 4 · Magnum v4