本文へスキップ
← ニュース一覧
Databricks Blog2026年8月6日

OfficeQA Pro V2 を発表:エンタープライズのグラウンデッド推論を測る新ベンチマーク

英語原文から AI が翻訳しました。 英語版を見る

要約

Databricks は、米国財務省の PDF 約 1,400 件から合成データパイプラインを用いて構築したエンタープライズ向けグラウンデッド推論ベンチマーク OfficeQA Pro V2 を公開した。ベンチマークの検証では

* OfficeQA Pro V2 is a new benchmark for enterprise grounded reasoning, built using our internal synthetic data pipeline from roughly 1,400 U.S. Treasury PDFs spanning 233 years and approximately 120,000 pages. * We leverage synthetic data generation to build OfficeQA Pro V2, allowing us to scale diverse, verified questions and answers. Combining these synthetic data techniques with our understanding of enterprise workflows enables us to rapidly build new benchmarks to make progress on the tasks our customers care about, like grounded reasoning. * An optimized agent harness can dramatically improve performance. Out-of-the-box agents averaged only 26.0% accuracy on OfficeQA Pro V2, while Databricks’ Genie delivered a 92% relative improvement on average across matched models and achieved up to 60% accuracy using the same models. Despite these gains, significant headroom remains in grounded reasoning.

関連記事

News

大規模環境における AI コーディングコストの管理

databricks-blog1d ago
News

AI アシスタントとは何か

databricks-blog1d ago
News

エージェント型ワークフローとは何か

databricks-blog2d ago
News

ツール呼び出し(tool calling)とは何か

databricks-blog2d ago