Models

Claude Opus 5 tops Artificial Analysis leaderboard

Artificial Analysis has updated its Intelligence Index to version 4.1.1, correcting grading errors and cementing Claude Opus 5's position at the top of the leaderboard with a score of 63.

AlphaSignal4 days agoModels
Image: AlphaSignal

Artificial Analysis has released version 4.1.1 of its Intelligence Index, a targeted patch designed to resolve grading discrepancies without altering benchmark weights or overall rankings. In this latest iteration, Claude Opus 5 retains its first-place position with an Intelligence Index score of 63. While most evaluated models experienced score adjustments of less than one point, Muse Spark 1.2 (xhigh) achieved the most significant upward shift, gaining 2.7 points due to enhanced grading robustness.

The update introduces two primary categories of corrections. First, the τ³-Banking benchmark has been upgraded to version 1.0.1, resolving scoring bugs associated with agent trajectories on multi-step banking tasks. Specifically, this fix addresses what the source calls an "unhappy path" where an AI agent successfully recovers from an error mid-task, alongside corrections to banking_knowledge task errors. Second, the evaluation platform has unified its grading infrastructure. The HLE, AA-LCR, and AA-Omniscience benchmarks now all utilize GPT-5.6 Luna as their grader model, replacing three separate legacy models to ensure evaluation consistency.

For AI practitioners and enterprise developers, this patch provides a more reliable framework for comparing model capabilities, particularly in complex, multi-step agentic workflows. The correction to the τ³-Banking benchmark means that previous performance metrics on this specific domain are not directly comparable to the new v1.0.1 results. However, the standardization under GPT-5.6 Luna reduces the noise introduced by disparate grading models. Developers can access these updated, standardized evaluations for free on the artificialanalysis.ai website, ensuring their deployment decisions are guided by highly consistent and robust benchmark data.

This is our own summary of reporting by AlphaSignal

More in Models