for agents/llms.txtv0.7.37 · 10 Oct 2026

Home / Articles / The infographic bake-off: which image model for which job, judged blind on 10 October 2026

The infographic bake-off: which image model for which job, judged blind on 10 October 2026

By · 2026-10-10 · article v1.0.0 · site v0.7.30 · infographicsimage-modelsopenroutergeminigpt-imageevaluationcostnewsroomdesigneraccountantdata-scientistarticle

Abstract: The articles here are moving from code-drawn figures to infographics made by image models, and the question was which model, for which job, at what cost. So we ran a bake-off. Every image-output model on OpenRouter on 10 October 2026, its auto-router and our own code-drawn renderer were given the same briefs from one article, from an eighteen-word stat card to the whole 2,800-word article, plus an edit, a UI component, a brand slide and a three-slide deck. Five rounds, models cut after each; 101 images, every one judged blind by a Designer agent; every cent recorded, $10.44 in all; an Accountant and a Data Scientist on the numbers. The result is three models and guidance. Nano Banana 2.1, at four cents an image, for concept slides, charts and decks. Gemini 3 Pro Image, at fourteen cents and twenty seconds, when you are in a hurry. GPT-5.4 Image 2, at seventeen to twenty-three cents and a minute and a half, for anything that starts from a document or where every character is specified. Choose the model by the job, not by the budget.

The bake-off's conclusion, drawn by the model it recommends for this job (GPT-5.4 Image 2, given the conclusion as a 263-word document). The blind judge scored it 9 out of 10 and would publish it as it is.

The articles on this site have carried figures drawn by code for months: diagrams from Mermaid, charts from SVG, decks from a renderer with fixed layouts. They are exact and they cost nothing, and they look like what they are. Some articles have also carried infographics made by an image model, one at a time, by hand. The next step is to make that part of the newsroom's workflow, which means an agent choosing a model, writing the brief and checking the result, every day.

That needs an answer to a plain question: which model, for which job, at what cost? So on 10 October we ran a bake-off, and this is what it found.

In short

The team

Four roles, each an agent with a role file. The Designer is adapted from the SG/Send team's Designer, whose first principle is that design is how it works: an infographic that is beautiful and wrong has failed. The judges saw copies of each image under a random id, never the model's name, and checked every required word and number before scoring anything. The SG/Send team has no accountant or data scientist, so both were written for this run. The Accountant's rule is that a failed image was still paid for, so the number that matters is cost per usable image, not price per image. The Data Scientist makes the cut after each round, and must not promote or drop a model on a difference smaller than the noise.

Round 1: who can put the right words in a picture?

Everyone got three short briefs. On the diagram, with five exact labels, the spread was wide:

The same diagram brief for all ten contestants, with the blind score. Two models ignored the requested 16:9 and returned a square that cut off boxes at both ends; one painted a muddy gradient; one misspelled three words.

Spelling turned out to be mostly solved: only one model, Gemini 2.5 Flash Image, misspelled labels. The faults were obedience and format. Most images added text nobody asked for: invented headings, "Source: internal analytics", page footers, a hex colour code from the prompt, "(16pt)" from a layout hint. GPT-5 Image and GPT-5 Image Mini ignore the requested aspect ratio and always return a 1024 square. The auto-router served Gemini 3 Pro Image, at its price. Three were cut.

Round 2: numbers, a document, a whole article

The chart had fourteen exact values. Five of six models drew every bar right, sorted and in proportion, with both highlight colours where asked:

Fourteen exact numbers. Five models got every value and proportion right; GPT-5 Image wrote "artitty" for artefact, three wrong values, three missing bars and a claim bar a third too short.

Then two briefs that start from text: the Explainer's 240-word two-minute version of the article as "here is a document, create an infographic about it", and the whole article. This is where the pattern I have seen with ChatGPT showed up, but only in the newest model: GPT-5.4 Image 2 scored 9 on both, and on the whole article it was the only model with every number right and the whole argument told. Give it the document, not a description of the picture. The older GPT-5 Image got worse with more text. The Google models summarised well but tended to drop the article's caveats ("re-anchoring does not enforce anything") and to invent plausible sample values.

The diagram brief ran again, unchanged, to measure noise: scores moved by up to two points between identical runs. One image per cell separates tiers, not neighbours.

Round 3: can you direct it?

Three briefs that test obedience. The hardest was an edit: send the model its own round 1 diagram and ask for three changes and nothing else.

The edit. GPT-5.4 Image 2 added a sixth box and kept the other five. Gemini 3 Pro changed the title and the colour, and made room for the new box by replacing the last one: "canary report" is gone.

Only GPT-5.4 Image 2 made all three changes without breaking anything (9). Every Google model treated the edit as a redraw, and something else moved: a box deleted, a label orphaned, a label overwritten. On the UI component (seven rows of exact text) and the brand slide, GPT-5.4 Image 2 again won both samples. Nano Banana 2.1, the best value of the first two rounds, had no usable image in this round: the briefs where every character is specified are the ones it loosens.

Round 4: three slides that belong together

Each model drew three slides of the same deck, each a separate call. The cheapest set was the best, Nano Banana 2.1's at 8, but no model carried the cover's top bar or footer to the other slides: nothing in a separate call knows what the last slide looked like.

So we tested the obvious fix straight away: send slide 1 with every later slide and ask for the same style. Judged blind in pairs, without knowing which set used the reference, every model's referenced set was the more consistent one: from 1 or 2 out of 5 to 4 or 5.

Nano Banana 2.1, slides 2 and 3 without and with slide 1 attached as a style reference. The top bar, the kicker and the headline type carry through.

Round 5: the conclusion, drawn by the models it recommends

The final three were given the conclusion as a 263-word document, written with the prompt rules the bake-off had produced. All six images got every tier, price and speed right; no LinkedIn logos, no leaked codes, no invented prices. GPT-5.4 Image 2 scored 9 and 9 and told the whole conclusion: it is the image at the top of this article.

What it cost

Quality against price for every model. The three recommended sit on the frontier; Gemini 3.1 Flash is beaten by Nano Banana 2.1 on both axes; GPT-5 Image is the most expensive and the worst.
ModelPrice per imageCost per usable imageMedian secondsMean score
Nano Banana 2.1$0.039$0.06696.4
Gemini 3 Pro Image$0.140$0.238206.6
GPT-5.4 Image 2$0.176$0.230838.0

Three things from the Accountant:

The guidance

The deliverable is a page any agent can follow, in the vault as analysis/guidance.md. The short version:

JobUseAvoid
Concept slide, data chartNano Banana 2.1the 1024-square models
Labelled diagram, stat cardGemini 3 Pro ImageNano Banana for exact labels
A deckNano Banana 2.1, with slide 1 attached to every later slideseparate calls with no reference
Document or whole article to infographicGPT-5.4 Image 2, given the documentGPT-5 Image
Edit an image, UI component, brand slideGPT-5.4 Image 2every Google model for edits
In a hurryGemini 3 Pro ImageGPT-5.4 Image 2 (83 seconds)

And the prompt rules the faults taught, which made round 5 clean:

  1. Never name the platform. "Suitable for LinkedIn" put LinkedIn logos on the image. Say "portrait 4:5".
  2. End with "No other text." Unrequested headings, sources and footers were the commonest fault.
  3. Say "no numbers other than those given". Every model invented plausible values somewhere.
  4. Keep hex codes, type sizes and separators out of text to be drawn. They get drawn.
  5. For an edit, say what must not change.
  6. For a deck, attach slide 1 to every later slide.
  7. Check every image. OCR and the judge agreed on 76% of images: OCR is a cheap alarm, not a judge.

What this does not show

What comes next

The guidance becomes the Storyteller's: the Five readers decks, drawn so far by code, get an image pass on Nano Banana 2.1 with slide 1 as the reference, and a document figure on GPT-5.4 Image 2, with the cost of each logged against the article. And the bake-off runs again whenever a new model appears, because the answer to "which model" has a date on it.


Written by agent@riskmandate.ai (Claude Opus 5.5, claude-opus-5-5) in the sgit.ai site session on 10 October 2026, from a voice memo by Dinis Cruz, who has editorial responsibility and paid for the run. Every image was generated through OpenRouter and judged blind by Claude agents under the role files in the vault; every cost is OpenRouter's recorded usage. AI-generated, disclosed as Article 50 of the EU AI Act asks.

Threads

Site & engineeringNews & evidence This article as a graph →

Builds on

All articles · All graphs

Want the next issue by email. One issue a week or so: what was published, what it adds up to, and what is worth your time. Subscribe to the SGit Newsroom →

← All articles