An AI dungeon master that lets you lose

A game you cannot fail is a slideshow.

Game mastering is improvisation under constraint. The table goes somewhere you did not prepare for, and you have thirty seconds to invent a village, staff it with three people who want different things, and remember that the party burned a bridge in this region two sessions ago. Language models are unusually well suited to that specific job — it is generation, memory and voice acting, which is most of what a GM does between combats.

What they are usually bad at is the other half: consequence. A safety-tuned assistant will not let your rogue actually die, will not let the villain's plan succeed, and will not describe a battlefield in terms that make it frightening. OpenRogue's long-context models — Euryale, Hermes 3 405B, Hermes 4, WizardLM-2 — will hold a campaign's history and let the dice mean something, whether you are running solo or prepping for a real table.

What AI dungeon master actually demands from a model

Improvisation with continuity

The party always goes off-map. The model has to invent on the spot and then remember what it invented four sessions later, which is a long-context problem more than a creativity one.

Real stakes

If failure is impossible, nothing that happens matters. A GM model has to be willing to let a plan fail, an NPC betray the party, and a bad roll cost something the players cared about.

NPCs who want things

The difference between a quest-giver and a character is that a character has an agenda that does not include the party. A good model gives NPCs motives that occasionally cut across yours.

Tone control

A grimdark siege campaign and a comic heist need completely different registers, and the model has to hold whichever one you set for hours rather than drifting toward generic high fantasy.

Prep material you can use at a real table

Statblock-shaped encounters, rumour tables, faction clocks, dungeon keys. Structured output that you can print or paste into your notes without rewriting it first.

Where filtered assistants fall down

The GM who will not let you fail

Filtered assistants steer toward positive outcomes. Ambushes turn out to be misunderstandings, the trap does not quite trigger, the villain monologues instead of acting. Play stops being a game and becomes a story the model is telling you about how well you are doing.

Bloodless combat

Combat described without violence is combat described without stakes. Mainstream models produce oddly abstract battles — 'the fight was intense' — because concrete description of injury sits close to what their training discourages.

Refusals on the interesting villains

Cults, torture-happy inquisitors, slavers, plague-bringers: the standard antagonists of the genre. A filtered model will hedge, sanitize, or decline to design the villain's actual plan, which leaves you running a campaign against an offscreen abstraction.

Session amnesia

Short context turns a campaign into a series of one-shots. The model forgets the NPC the party adopted, invents a different name for the same town, and quietly contradicts a plot thread you had been building for a month.

Which model for AI dungeon master?

Models suited to AI dungeon master
ModelRoleWhy
EuryaleThe session-runnerIt writes with colour while keeping plot logic intact, which is precisely the GM balance, and 131K context carries a long campaign's history. It also handles several NPCs in one scene without their voices merging, which is where most models fail at a table.
Hermes 3 405BCampaign architectureUse it between sessions for the heavy design work: faction agendas, a villain's plan that survives contact with clever players, and the causal chain that connects a dungeon to a war. Its depth is worth the slower replies when you are prepping rather than playing.
RocinanteFast solo playFor rapid turn-by-turn solo sessions, near-instant replies matter more than depth — waiting on a large model breaks the rhythm of play. Rocinante writes well above its size and is deliberately game for wherever a scene goes.
WizardLM-2Prep documents and tablesMulti-constraint structured output is its specialty: a twelve-room dungeon key with a hazard, a clue and a faction presence per room, produced in one pass without losing the format. It is the model to use for material you will print.

A workflow that works

  1. Put the campaign bible in a Project: System and tone, the party roster, the map, the factions, and the running log of what has happened. Everything the model needs to not contradict itself lives here, and it is visible on every turn.
  2. State the rules of engagement: Tell it explicitly: you narrate outcomes, I roll; failure is allowed; never decide what my character does or feels; end each turn with a decision point. Most bad AI-GM experiences are a missing instruction rather than a bad model.
  3. Prep between sessions with the big model: Hermes 3 405B for the villain's plan and the faction clocks, WizardLM-2 for the dungeon key and the rumour table. Paste the results back into the campaign bible so the play session has them in context.
  4. Run the session on a faster model: Switch to Euryale for a full table-style session, or Rocinante for quick solo play. Speed is a feature during play in a way it never is during prep.
  5. Keep the dice yours: Roll physically or with a dice app and tell the model the result. Models are unreliable random number generators and will unconsciously favour outcomes that suit the story — which is the one place you do not want a collaborator.
  6. Log the session before you close it: Ask for a five-bullet recap: what changed, who is owed what, what is still unresolved. Append it to the campaign bible. That log is what makes session twenty feel connected to session three.

Where this approach has limits

This is a chat window, not a virtual tabletop. There is no dice roller, no character sheet, no initiative tracker, no map and no integration with any VTT — you bring those. Models are also unreliable rules lawyers: they will state a rule from a game system with total confidence and be wrong, blending editions and inventing interactions, so anything mechanical needs checking against the actual rulebook. And an AI GM is genuinely not a substitute for a human one at a real table; it is very good for solo play, for prep, and for improvising an NPC you did not plan, and it does not do the thing a good human GM does when they read the room.

Prompts to start from

/// Rules of play
You narrate, I roll. Failure is allowed. End every turn on a decision.

/// Prep
Design a villain whose plan still works if the party wins the first two fights.

/// Table
Twelve-room dungeon key: hazard, clue and faction presence for each room.

Frequently asked questions

Can an AI actually run a tabletop campaign?

For solo play, yes — genuinely well, if you set the rules of engagement in your first message and keep the campaign history in a Project. For a real table, it is better used as a prep and improvisation tool: encounter design, NPC voices, and the village you did not plan for. It does not read a room the way a human GM does.

Which model makes the best AI dungeon master?

Euryale 70B for running sessions — expressive, coherent, and good at keeping multiple NPCs distinct, with 131K context for campaign memory. Hermes 3 405B for between-session design work, and WizardLM-2 8x22B for structured prep material like dungeon keys and rumour tables.

Will it let my character die?

It will if you tell it to. Uncensored models have no trained reluctance to narrate failure or serious consequence, but assistant instincts still push toward accommodation, so state up front that failure is allowed and that you want outcomes narrated honestly. Once that instruction is in the Project, it holds.

Does it know the rules of my game system?

Approximately, and not reliably enough to trust. Models blend editions, invent plausible-sounding interactions and state wrong rules confidently. Use it for fiction, NPCs and encounter design; check anything mechanical against the rulebook, and roll your own dice.

How do I keep a long campaign consistent?

Keep a campaign bible in your Project's custom instructions and append a short recap after every session. The 131K-context models can then see the whole history alongside the current scene, which is what prevents the model from renaming a town or forgetting the NPC your party adopted.

Related

Start free → · All models · Pricing