Databricks Launches OfficeQA Pro V2 AI Benchmark
Databricks has released OfficeQA Pro V2, a challenging new benchmark designed to evaluate how well AI agents perform complex grounded reasoning across massive, messy enterprise document collections.

Databricks has released OfficeQA Pro V2, a benchmark evaluating AI agent generalization on enterprise grounded reasoning. Built using the asynth synthetic data library, it features 90 questions grounded in 1,400 U.S. Treasury PDFs spanning 233 years (1793 to 2024) and 120,000 pages, released for the nation's 250th anniversary. While the original OfficeQA, created in late 2025 with 89,000 pages, averaged 2 source documents per question for its OfficeQA Pro subset, V2 averages 6.7 (median 5.5, max 24). In V2, 74.4% of questions require four or more sources, compared to 62.4% in OfficeQA Pro and 56.1% in OfficeQA Full. Additionally, 7% of questions require visual understanding (up from 3% in OfficeQA Pro), and 10% require web search (compared to 21.8% in OfficeQA Pro and 15.9% in OfficeQA Full).
The benchmark, originally used for the Grounded Reasoning Cup where 11 academic teams averaged 41.1% accuracy and the winner hit 63.3%, remains difficult. Out-of-the-box baseline agents using Claude Code or Codex averaged 26.0% accuracy across five models. Frontier agents using Claude Opus 4.8, Claude Fable 5, GPT-5.5, or GPT-5.6 Sol averaged 37.5%. Specifically, Sonnet 5 on Claude Code scored 15.6% at $5.01 per rollout, while GPT-5.6 Sol on Codex scored 33.3% at $4.70.
Databricks Genie, utilizing ai_parse, significantly boosted performance. Across matched models, Genie improved accuracy by an average of 24.0 percentage points (a 92% relative improvement), raising the four-model mean by 15.3 percentage points (from 37.5% to 52.8%) and reaching up to 60% accuracy. Genie configurations using GPT-5.6 Luna, GPT-5.6 Terra, and Claude Fable 5 dominated the cost-quality frontier. For Claude Fable 5, Genie improved accuracy by 14.4 percentage points (+32% relative) while cutting costs 9x from Claude Code's looping-heavy $37.36 average.
For practitioners, OfficeQA Pro V2—available on Hugging Face with evaluation code on GitHub—proves that optimized agent harnesses are critical. It shows that while raw models struggle with historical shifts across 232 years of financial reporting, structured parsing and retrieval can unlock massive efficiency and accuracy gains.
This is our own summary of reporting by Databricks AI



