No single model wins at text-to-image.
We blind tested eight image models with human graders, then built a router that picks the right one for the job.

There is no one-size-fits-all image model.
For the work we do at Frontward we need the best model across a wide range of image tasks, and the best one keeps changing depending on the task. Unlike coding models, the output here is subjective. There is no verifiable truth to check against.
That made it an interesting problem. We decided to build a benchmark on human judgement, and to use it to guide a model router that picks the right image model for the job.
Blind tests, human graders.
We broke our work down into seven categories: logos, illustration, photorealism, text rendering, product shots, badges, and posters. Then we took what we considered the eight best image models and blind tested them.
Graders were shown two images from different models given the same prompt, names hidden, and picked the one that matched the prompt more accurately. Ratings are Elo, starting at 1000 and replayed in the order the votes came in. There is no automatic judge anywhere. A rating only moves when a person picks.
Speed and cost are averaged separately over each model's runs, because a model that wins in forty seconds is a different answer from one that wins in ten.
Eight models, ranked by blind votes.
Elo and win rate come from the votes; speed and cost are averaged over each model's runs.
- 01
GPT Image 2 OpenAI
1108
64.9%
40.5s
$0.05
- 02
Nano Banana Pro Google
1083
61.1%
22.7s
$0.02
- 03
FLUX.2 Black Forest Labs
1037
63.6%
9.8s
$0.04
- 04
Nano Banana 2 Google
1009
48.4%
8.3s
$0.07
- 05
Recraft V4.1 Recraft
1008
45.0%
6.8s
$0.04
- 06
FLUX.2 Max Black Forest Labs
1004
50.9%
23.7s
$0.09
- 07
Grok Imagine xAI
918
33.3%
6.8s
$0.02
- 08
Recraft V3 Recraft
834
10.5%
5.9s
$0.04
The best model is the slowest one.
Every model on the board, plotted by Elo against how long it takes. Bigger bubble, more expensive.
Every model, every category.
The overall board hides most of the story. Each model's Elo within each category, brighter is better.
No one model does everything well.
GPT Image 2 is probably the best model overall. It is also the slowest by a wide margin. Fine for a hero image, wrong for anything a user is sitting and waiting on.
Nano Banana Pro is great at photorealism, product shots, and text rendering, and it is one of the cheapest models in the set. GPT Image 2, for all its overall lead, lands near the bottom on text. The FLUX models were incredible at creative work: illustration, design, and branding. FLUX.2 wins illustration outright, by 77 Elo on the best-sampled category we have, and comes back in under ten seconds at four cents.
Grok was garbage at pretty much every category except logos, and logos was our least tested category, so take that with some salt.
The surprising part: price and quality barely correlate. That is not true for coding models. The second-best model here is nearly the cheapest, and the most expensive one lands mid-pack.
Fine-tuning still beats all of this.
There is a ton of nuance behind the table. The middle of the board is close, and prices are what providers charged in August 2026, which will move.
Realistically, a brand with the skill set and resources should fine-tune with LoRAs. That has by far produced the best results for us, even on small local base models, compared to GPT Image 2 out of the box. That probably deserves a post of its own.
A small model router built on the results.
All of this went into a small router. It reads the prompt, sorts it into a category, picks a frame from the wording (a banner wants 16:9, a story wants 9:16), and sends the request to that category's winner.
The routing table is the benchmark, not our opinion. Each row cites its standing, carries a confidence (strong, weak, or untested), and every pick returns a one-line reason.
When we re-run the benchmark and the standings move, the router moves with them. A row only changes when the evidence does.
Building something with image models?
We route, evaluate, and ship them in production for clients.