Aller au contenu
← Toutes les actus
Databricks Blog6 août 2026

OfficeQA Pro V2 : un nouveau benchmark pour le raisonnement ancré en entreprise

Traduit de l'original anglais par IA. Voir en anglais

Résumé

Databricks publie OfficeQA Pro V2, un nouveau benchmark de raisonnement ancré en entreprise, construit à partir d'environ 1 400 PDF du Trésor américain via un pipeline de données synthétiques. Les tests du benchmark

* OfficeQA Pro V2 is a new benchmark for enterprise grounded reasoning, built using our internal synthetic data pipeline from roughly 1,400 U.S. Treasury PDFs spanning 233 years and approximately 120,000 pages. * We leverage synthetic data generation to build OfficeQA Pro V2, allowing us to scale diverse, verified questions and answers. Combining these synthetic data techniques with our understanding of enterprise workflows enables us to rapidly build new benchmarks to make progress on the tasks our customers care about, like grounded reasoning. * An optimized agent harness can dramatically improve performance. Out-of-the-box agents averaged only 26.0% accuracy on OfficeQA Pro V2, while Databricks’ Genie delivered a 92% relative improvement on average across matched models and achieved up to 60% accuracy using the same models. Despite these gains, significant headroom remains in grounded reasoning.

Articles similaires

News

Comment j'ai conçu des revues de sécurité basées sur des agents sur Databricks

databricks-blog11h ago
News

Comment Concurrence gouverne l'IA clinique à l'échelle du billion de tokens avec Unity Gateway

databricks-blog1d ago
News

Le Genie One MCP est désormais généralement disponible (GA)

databricks-blog1d ago
News

Genie One MCP : offrez à n'importe quel agent IA le bon contexte métier

databricks-blog1d ago