Blog · August 26, 2026

No single model wins at text-to-image.

We blind tested eight image models with human graders, then built a router that picks the right one for the job.

Painting of a figure under a lamp comparing eight canvases leaning against a studio wall
Why

There is no one-size-fits-all image model.

For the work we do at Frontward we need the best model across a wide range of image tasks, and the best one keeps changing depending on the task. Unlike coding models, the output here is subjective. There is no verifiable truth to check against.

That made it an interesting problem. We decided to build a benchmark on human judgement, and to use it to guide a model router that picks the right image model for the job.

Method

Blind tests, human graders.

We broke our work down into seven categories: logos, illustration, photorealism, text rendering, product shots, badges, and posters. Then we took what we considered the eight best image models and blind tested them.

Graders were shown two images from different models given the same prompt, names hidden, and picked the one that matched the prompt more accurately. Ratings are Elo, starting at 1000 and replayed in the order the votes came in. There is no automatic judge anywhere. A rating only moves when a person picks.

Speed and cost are averaged separately over each model's runs, because a model that wins in forty seconds is a different answer from one that wins in ten.

Standings

Eight models, ranked by blind votes.

Elo and win rate come from the votes; speed and cost are averaged over each model's runs.

  1. 01

    GPT Image 2 OpenAI

    1108

    64.9%

    40.5s

    $0.05

  2. 02

    Nano Banana Pro Google

    1083

    61.1%

    22.7s

    $0.02

  3. 03

    FLUX.2 Black Forest Labs

    1037

    63.6%

    9.8s

    $0.04

  4. 04

    Nano Banana 2 Google

    1009

    48.4%

    8.3s

    $0.07

  5. 05

    Recraft V4.1 Recraft

    1008

    45.0%

    6.8s

    $0.04

  6. 06

    FLUX.2 Max Black Forest Labs

    1004

    50.9%

    23.7s

    $0.09

  7. 07

    Grok Imagine xAI

    918

    33.3%

    6.8s

    $0.02

  8. 08

    Recraft V3 Recraft

    834

    10.5%

    5.9s

    $0.04

Quality vs speed

The best model is the slowest one.

Every model on the board, plotted by Elo against how long it takes. Bigger bubble, more expensive.

800900100011005s10s20s40sSeconds per image, log scale → slowerElo → betterGPT Image 2$0.05Nano Banana Pro$0.02FLUX.2$0.04Nano Banana 2$0.07Recraft V4.1$0.04FLUX.2 Max$0.09Grok Imagine$0.02Recraft V3$0.04Cost per image$0.02$0.05$0.09
Elo from blind human votes. Speed and cost averaged per run. Bubble area is cost per image.
By category

Every model, every category.

The overall board hides most of the story. Each model's Elo within each category, brighter is better.

Logos
Illustration
Photo
Text
Product
Badges
Posters
GPT Image 2 1108
990
1059
1011
971
1036
1071
1082
Nano Banana Pro 1083
1027
938
1046
1061
1056
996
1039
FLUX.2 1037
970
1136
1012
1002
1031
1003
983
Nano Banana 2 1009
1007
982
973
1039
971
1050
940
Recraft V4.1 1008
1033
1002
968
984
995
987
1004
FLUX.2 Max 1004
970
1058
1035
998
975
983
1012
Grok Imagine 918
1034
847
986
962
999
926
999
Recraft V3 834
969
979
968
984
937
986
940
Category Elo from blind pairwise votes, August 2026. Brighter is better; the outlined cell won its category.
What we found

No one model does everything well.

GPT Image 2 is probably the best model overall. It is also the slowest by a wide margin. Fine for a hero image, wrong for anything a user is sitting and waiting on.

Nano Banana Pro is great at photorealism, product shots, and text rendering, and it is one of the cheapest models in the set. GPT Image 2, for all its overall lead, lands near the bottom on text. The FLUX models were incredible at creative work: illustration, design, and branding. FLUX.2 wins illustration outright, by 77 Elo on the best-sampled category we have, and comes back in under ten seconds at four cents.

Grok was garbage at pretty much every category except logos, and logos was our least tested category, so take that with some salt.

The surprising part: price and quality barely correlate. That is not true for coding models. The second-best model here is nearly the cheapest, and the most expensive one lands mid-pack.

The nuance

Fine-tuning still beats all of this.

There is a ton of nuance behind the table. The middle of the board is close, and prices are what providers charged in August 2026, which will move.

Realistically, a brand with the skill set and resources should fine-tune with LoRAs. That has by far produced the best results for us, even on small local base models, compared to GPT Image 2 out of the box. That probably deserves a post of its own.

The router

A small model router built on the results.

All of this went into a small router. It reads the prompt, sorts it into a category, picks a frame from the wording (a banner wants 16:9, a story wants 9:16), and sends the request to that category's winner.

The routing table is the benchmark, not our opinion. Each row cites its standing, carries a confidence (strong, weak, or untested), and every pick returns a one-line reason.

When we re-run the benchmark and the standings move, the router moves with them. A row only changes when the evidence does.

Work with us

Building something with image models?

We route, evaluate, and ship them in production for clients.