A million tokens of context, at Flash speed.
Google DeepMind · Gemini Flash · Not disclosed parameters · 1M context · Pro · $20/mo · Filtered by its provider
The Flash line is Google's answer to a specific question: what if the enormous context window were also the cheap, fast option? Gemini 3.5 Flash carries a million-token window at latency that feels closer to a small model than a frontier one, and it reads images, PDFs and long structured documents natively. That combination makes it the bulk-work model in the OpenRogue lineup. Drop in an entire archive — a year of support tickets, a full technical manual, every chapter you have written so far, a folder of scanned paperwork — and ask questions across the whole thing without chunking it, indexing it, or waiting a minute per answer. For extraction, cross-referencing and 'find me every instance of' work, nothing else in the dropdown is as well suited.
Google's caution comes with it. Gemini models are wrapped in safety classifiers tuned by a company that answers a billion queries a day and has no appetite for being quoted saying something interesting, so anything touching politics, medicine, sexuality or violence tends to come back hedged, disclaimed or declined. Ask it to summarise a thousand pages and it is superb; ask it what it thinks about the argument inside them and it will find a way to say nothing. So use it for the mechanical half of the work and switch when you get to the opinionated half — Hermes 4 70B for analysis it will actually commit to, Hermes 3 405B when the judgement call is genuinely hard. The conversation carries over.
No. Google applies safety classifiers and policy training across its consumer AI models, and anything touching politics, health specifics, sexuality or violence tends to be hedged or refused. OpenRogue adds nothing on top of that, and cannot remove it either — which is why the uncensored models sit one click away in the same dropdown.
Volume. The million-token window plus low latency makes it the right model for reading an entire archive at once — extraction, cross-referencing, deduplication, building a table out of three hundred documents. It is the model you use when the hard part is the size of the input rather than the difficulty of the question.
Both carry a million-token window. Flash is faster and better suited to mechanical passes over huge inputs; Sonnet is the better reader when the task needs judgement, nuance or writing quality on the way out. They cost different amounts of compute, so it is worth trying the cheap one first.
Hermes 4 70B is the natural next step — it reasons deliberately, has a 131K context window, and Nous Research trains it for neutral alignment rather than corporate policy. For a genuinely hard judgement call, Hermes 3 405B goes deeper. For creative work, Magnum v4 72B or Euryale 70B.
The model family is the same; the product is not. Google's assistant is woven into Gmail, Docs and Android and carries additional product-level policy. Here it is one option in a dropdown alongside genuinely uncensored models, with no Google integration and no extra filter from us. The comparison page has the full breakdown.
Google does not publish parameter counts for the Gemini line, so we do not quote one. What is published and what actually matters in use is the context window and the multimodal input — a million tokens, with images and documents read natively.
Related models: Hermes 4 · Hermes 3 405B · GPT-4o