Google Gemini 3.6 Flash hits 91.2% on reasoning test
Google's Gemini 3.6 Flash has scored 91.2% on the ARC-AGI-1 reasoning test, showcasing a highly competitive balance of cost and fluid intelligence for AI developers.

Google has officially entered two new models, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, into the ARC Prize verified leaderboard. The standout performer, Gemini 3.6 Flash, achieved a 91.2% score on the ARC-AGI-1 benchmark at a cost of $0.34 per task. On the more challenging ARC-AGI-2 benchmark, the model scored 60.4% at $0.61 per task. Meanwhile, the budget-friendly Gemini 3.5 Flash-Lite scored just 10.3% on ARC-AGI-2, though it cost nearly four times less at $0.14 per task.
Created by François Chollet, the ARC-AGI benchmark evaluates fluid intelligence by requiring models to deduce grid transformation rules from colored square examples with zero partial credit and a maximum of three attempts. The evaluation tested both models across four reasoning effort levels: High, Medium, Low, and Minimal. The impact of compute allocation was stark; Gemini 3.6 Flash's performance on ARC-AGI-2 plummeted from its peak of 60.4% at High effort down to a mere 2.6% at Minimal effort.
These results place Gemini 3.6 Flash in a competitive mid-tier on the ARC-AGI-2 leaderboard. It still trails top-tier models like GPT-5.6 Sol at 92.5%, Claude Opus 5 at 90.4%, and GPT-5.5 at 85%, the latter of which meets the grand prize threshold of over 85%. For comparison, an individual human averages 66% on the test, while a human panel reaches 100%.
For AI practitioners, these public results on arcprize.org highlight a critical trade-off between cost and cognitive depth. While Gemini 3.5 Flash-Lite offers extreme affordability, its low reasoning score makes it unsuitable for complex, novel problem-solving. Conversely, Gemini 3.6 Flash proves that mid-tier models can approach human-level fluid intelligence when allowed sufficient reasoning compute, offering a viable, cost-effective alternative for tasks requiring adaptable logic rather than rote memorization.
This is our own summary of reporting by AlphaSignal



