Qwen3.8 Max tops Artificial Analysis agentic index
Qwen3.8 Max has secured the top spot on Artificial Analysis's agentic index, marking a significant shift in the competitive landscape for autonomous AI workflows.

Artificial Analysis has released version 4.1.1 of its Intelligence Index, revealing major shifts in model rankings, including Qwen3.8 Max claiming the top spot on the agentic index. The update also introduces Claude Opus 5, specifically the Adaptive Reasoning, Low Effort variant, which matches the intelligence of Fable 5 while reducing the financial expense per task. Meanwhile, DeepSeek V4 Flash 0731, utilizing Reasoning, Max Effort, scored a 50 on the Intelligence Index, representing a 10-point increase over the previous DeepSeek V4 Flash. Additionally, Inkling Small achieved a score within a single point of the standard Inkling model despite operating with less than a third of its parameters.
The v4.1.1 update incorporates nine distinct evaluations: GDPval-AA v2, τ³-Banking (now moved to v1.0.1), Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. To refine these tests, the grading system for Humanity's Last Exam, AA-LCR, and AA-Omniscience has been upgraded to GPT-5.6 Luna (medium). The platform also updated its Cost per Task methodology, resulting in minor absolute increases in cost estimates while preserving relative model positions. Other evaluated benchmarks include the AA-Briefcase agentic evaluation, which tests realistic business workflows, and coding agent benchmarks like DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA.
For AI practitioners, these updates provide a highly granular framework for choosing models based on specific trade-offs between intelligence, speed, and cost. The introduction of the Endpoint Accuracy Index, which runs BFCL v4-500, HLE-250, and AA-LCR-25 against various provider endpoints, helps developers determine if third-party hosts like Groq, Cerebras, Fireworks, or DeepInfra preserve the accuracy of reference models like gpt-oss-120b (high). With agentic workflows demanding high autonomy, the rise of specialized reasoning models like Qwen3.8 Max and Claude Opus 5 allows developers to deploy highly capable agents for complex problem-solving while managing token costs more predictably.
This is our own summary of reporting by Hacker News



