Artificial Analysis Revamps Intelligence Index v5
Independent evaluator Artificial Analysis will release its Intelligence Index v5 in late October, introducing private coding tests and scientific terminal tasks to combat benchmark contamination.

Artificial Analysis is preparing to launch Intelligence Index v5 in late October, marking its most significant benchmark overhaul of the year. The update introduces Terminal-Bench Science and a new, unpublished coding evaluation designed to prevent data contamination. Because researchers frequently cite this index, the upcoming changes will alter model rankings and prevent direct comparisons with older scores. The current v4.3.2 version relies on category weights of 30% for Agents, 20% for Coding, 30% for General, and 20% for Scientific Reasoning.
The current v4.3.2 build aggregates several evaluations, including AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. Version 5 will expand this list. Its new Terminal-Bench Science component, developed by Stanford University researchers alongside the Terminal-Bench and Harbor teams, tests whether AI agents can execute realistic scientific workflows in a command-line environment. This benchmark features 70 tasks across five domains, with each task attempted three times and scored using a strict pass@1 metric.
Early leaderboards for Terminal-Bench Science show that current systems still struggle with these complex workflows. GPT-6 Astra Max leads the evaluation at 63.3%, followed by Claude Opus 5.5 (Xhigh, Default Fallback) at 61.9% and Claude Opus 5.5 (Max, Default Fallback) at 59.0%. To address the risk of models memorizing public training data, Artificial Analysis is also adding a private coding benchmark with undisclosed tasks and datasets, building on its previous integration of AutomationBench-AA.
For practitioners, these updates require adjusting how they evaluate model performance. Developers should treat v4.3.2 and v5 scores as separate series until the final weights are published. Additionally, teams must update their coding-agent baselines, as Terminal-Bench 4.0 replaces v2.1 with a more difficult 66-task set featuring revised environments and verifiers. Finally, tracking task-level failures across planning, shell use, and dependency management will be crucial for diagnosing why research agents fail.
This is our own summary of reporting by AlphaSignal


