A game you cannot fail is a slideshow.
Game mastering is improvisation under constraint. The table goes somewhere you did not prepare for, and you have thirty seconds to invent a village, staff it with three people who want different things, and remember that the party burned a bridge in this region two sessions ago. Language models are unusually well suited to that specific job — it is generation, memory and voice acting, which is most of what a GM does between combats.
What they are usually bad at is the other half: consequence. A safety-tuned assistant will not let your rogue actually die, will not let the villain's plan succeed, and will not describe a battlefield in terms that make it frightening. OpenRogue's long-context models — Euryale, Hermes 3 405B, Hermes 4, WizardLM-2 — will hold a campaign's history and let the dice mean something, whether you are running solo or prepping for a real table.
The party always goes off-map. The model has to invent on the spot and then remember what it invented four sessions later, which is a long-context problem more than a creativity one.
If failure is impossible, nothing that happens matters. A GM model has to be willing to let a plan fail, an NPC betray the party, and a bad roll cost something the players cared about.
The difference between a quest-giver and a character is that a character has an agenda that does not include the party. A good model gives NPCs motives that occasionally cut across yours.
A grimdark siege campaign and a comic heist need completely different registers, and the model has to hold whichever one you set for hours rather than drifting toward generic high fantasy.
Statblock-shaped encounters, rumour tables, faction clocks, dungeon keys. Structured output that you can print or paste into your notes without rewriting it first.
Filtered assistants steer toward positive outcomes. Ambushes turn out to be misunderstandings, the trap does not quite trigger, the villain monologues instead of acting. Play stops being a game and becomes a story the model is telling you about how well you are doing.
Combat described without violence is combat described without stakes. Mainstream models produce oddly abstract battles — 'the fight was intense' — because concrete description of injury sits close to what their training discourages.
Cults, torture-happy inquisitors, slavers, plague-bringers: the standard antagonists of the genre. A filtered model will hedge, sanitize, or decline to design the villain's actual plan, which leaves you running a campaign against an offscreen abstraction.
Short context turns a campaign into a series of one-shots. The model forgets the NPC the party adopted, invents a different name for the same town, and quietly contradicts a plot thread you had been building for a month.
| Model | Role | Why |
|---|---|---|
| Euryale | The session-runner | It writes with colour while keeping plot logic intact, which is precisely the GM balance, and 131K context carries a long campaign's history. It also handles several NPCs in one scene without their voices merging, which is where most models fail at a table. |
| Hermes 3 405B | Campaign architecture | Use it between sessions for the heavy design work: faction agendas, a villain's plan that survives contact with clever players, and the causal chain that connects a dungeon to a war. Its depth is worth the slower replies when you are prepping rather than playing. |
| Rocinante | Fast solo play | For rapid turn-by-turn solo sessions, near-instant replies matter more than depth — waiting on a large model breaks the rhythm of play. Rocinante writes well above its size and is deliberately game for wherever a scene goes. |
| WizardLM-2 | Prep documents and tables | Multi-constraint structured output is its specialty: a twelve-room dungeon key with a hazard, a clue and a faction presence per room, produced in one pass without losing the format. It is the model to use for material you will print. |
This is a chat window, not a virtual tabletop. There is no dice roller, no character sheet, no initiative tracker, no map and no integration with any VTT — you bring those. Models are also unreliable rules lawyers: they will state a rule from a game system with total confidence and be wrong, blending editions and inventing interactions, so anything mechanical needs checking against the actual rulebook. And an AI GM is genuinely not a substitute for a human one at a real table; it is very good for solo play, for prep, and for improvising an NPC you did not plan, and it does not do the thing a good human GM does when they read the room.
/// Rules of play
You narrate, I roll. Failure is allowed. End every turn on a decision.
/// Prep
Design a villain whose plan still works if the party wins the first two fights.
/// Table
Twelve-room dungeon key: hazard, clue and faction presence for each room.
For solo play, yes — genuinely well, if you set the rules of engagement in your first message and keep the campaign history in a Project. For a real table, it is better used as a prep and improvisation tool: encounter design, NPC voices, and the village you did not plan for. It does not read a room the way a human GM does.
Euryale 70B for running sessions — expressive, coherent, and good at keeping multiple NPCs distinct, with 131K context for campaign memory. Hermes 3 405B for between-session design work, and WizardLM-2 8x22B for structured prep material like dungeon keys and rumour tables.
It will if you tell it to. Uncensored models have no trained reluctance to narrate failure or serious consequence, but assistant instincts still push toward accommodation, so state up front that failure is allowed and that you want outcomes narrated honestly. Once that instruction is in the Project, it holds.
Approximately, and not reliably enough to trust. Models blend editions, invent plausible-sounding interactions and state wrong rules confidently. Use it for fiction, NPCs and encounter design; check anything mechanical against the rulebook, and roll your own dice.
Keep a campaign bible in your Project's custom instructions and append a short recap after every session. The 131K-context models can then see the whole history alongside the current scene, which is what prevents the model from renaming a town or forgetting the NPC your party adopted.