Zum Inhalt springen
← Alle News
Databricks Blog6. August 2026

OfficeQA Pro V2: ein neuer Benchmark für Grounded Reasoning im Unternehmen

Von KI aus dem englischen Original übersetzt. Auf Englisch ansehen

Zusammenfassung

Databricks hat OfficeQA Pro V2 veröffentlicht, einen neuen Benchmark für Grounded Reasoning im Unternehmenskontext, der aus rund 1.400 PDFs des US-Finanzministeriums über eine Pipeline für synthetische Daten aufgebaut wurde. Die Benchmark-Tests

* OfficeQA Pro V2 is a new benchmark for enterprise grounded reasoning, built using our internal synthetic data pipeline from roughly 1,400 U.S. Treasury PDFs spanning 233 years and approximately 120,000 pages. * We leverage synthetic data generation to build OfficeQA Pro V2, allowing us to scale diverse, verified questions and answers. Combining these synthetic data techniques with our understanding of enterprise workflows enables us to rapidly build new benchmarks to make progress on the tasks our customers care about, like grounded reasoning. * An optimized agent harness can dramatically improve performance. Out-of-the-box agents averaged only 26.0% accuracy on OfficeQA Pro V2, while Databricks’ Genie delivered a 92% relative improvement on average across matched models and achieved up to 60% accuracy using the same models. Despite these gains, significant headroom remains in grounded reasoning.

Ähnliche Artikel

News

KI-Coding-Kosten im großen Maßstab steuern

databricks-blog1d ago
News

Was ist ein KI-Assistent?

databricks-blog1d ago
News

Was sind Agentic Workflows?

databricks-blog2d ago
News

Was ist Tool Calling?

databricks-blog2d ago