Research

Artificial Analysis Overhauls Image Arena to Rank Models

Artificial Analysis has overhauled its Text to Image Arena to rank models across 10 distinct use cases, helping developers choose the best model for their specific production needs.

AlphaSignal3 days agoResearch
Image: AlphaSignal

The AI evaluation platform Artificial Analysis has redesigned its Text to Image Arena, moving away from a single aggregate score to rank models across 10 real-world use cases and nine capabilities. To prevent leaderboard overfitting, the platform now refreshes its evaluation prompts monthly using anonymized data from real users, retiring prompts once they no longer help differentiate between models. The system relies on blind human preference testing to calculate Elo ratings, utilizing judging hints to guide voters on complex prompts.

In the updated rankings, GPT Image 2 leads every sub-leaderboard with an Elo rating of 1339. However, this top-tier performance comes at a high price of $211 per 1,000 images. Holding the second-place spot overall is Reve 2.1 with an Elo rating of 1299. Developed by a small lab of approximately 65 people and trained on ten times fewer GPUs than its competitors, this 4K layout-first model manages to tie GPT Image 2 specifically on user interface and user experience tasks.

For budget-conscious practitioners, Nano Banana 2 emerges as the primary value alternative. Priced at $67 per 1,000 images, which is roughly one-third the cost of GPT Image 2, it serves as a strong choice for UI/UX, text rendering, and animation and gaming use cases. Across all top-10 models evaluated on the platform, understanding physics remains the weakest capability overall, with no model ranking it as its strongest suite.

This granular benchmarking shift means developers and product teams no longer have to rely on generic quality scores that fail to reflect specialized workflows. A game studio designing concept art, a marketing agency generating campaign assets, or a software team prototyping UI mockups can now isolate the exact capabilities they need. By matching specific project requirements to targeted model strengths, teams can optimize both performance and API costs.

This is our own summary of reporting by AlphaSignal

More in Research