Pular para o conteúdo
← Todas as notícias
Databricks Blog6 de agosto de 2026

Apresentando o OfficeQA Pro V2: um novo benchmark para raciocínio fundamentado no contexto corporativo

Traduzido do original em inglês por IA. Ver em inglês

Resumo

A Databricks lançou o OfficeQA Pro V2, um novo benchmark de raciocínio fundamentado para o contexto corporativo, construído a partir de cerca de 1.400 PDFs do Tesouro dos EUA com um pipeline de dados sintéticos. Os testes do benchmark

* OfficeQA Pro V2 is a new benchmark for enterprise grounded reasoning, built using our internal synthetic data pipeline from roughly 1,400 U.S. Treasury PDFs spanning 233 years and approximately 120,000 pages. * We leverage synthetic data generation to build OfficeQA Pro V2, allowing us to scale diverse, verified questions and answers. Combining these synthetic data techniques with our understanding of enterprise workflows enables us to rapidly build new benchmarks to make progress on the tasks our customers care about, like grounded reasoning. * An optimized agent harness can dramatically improve performance. Out-of-the-box agents averaged only 26.0% accuracy on OfficeQA Pro V2, while Databricks’ Genie delivered a 92% relative improvement on average across matched models and achieved up to 60% accuracy using the same models. Despite these gains, significant headroom remains in grounded reasoning.

Artigos relacionados

News

Como criei revisões de segurança baseadas em agentes no Databricks

databricks-blog11h ago
News

Como a Concurrence governa a IA clínica em escala de trilhões de tokens com o Unity Gateway

databricks-blog1d ago
News

O Genie One MCP agora está geralmente disponível (GA)

databricks-blog1d ago
News

Genie One MCP: forneça o contexto de negócios correto para qualquer agente de IA

databricks-blog1d ago