OfficeQA Pro V2 공개: 기업용 근거 기반 추론을 위한 새로운 벤치마크
영어 원문을 AI가 번역했습니다. 영어로 보기
Databricks가 미국 재무부 PDF 약 1,400건을 합성 데이터 파이프라인으로 가공해 만든 기업용 근거 기반 추론 벤치마크 OfficeQA Pro V2를 공개했다. 벤치마크 테스트에서는
* OfficeQA Pro V2 is a new benchmark for enterprise grounded reasoning, built using our internal synthetic data pipeline from roughly 1,400 U.S. Treasury PDFs spanning 233 years and approximately 120,000 pages. * We leverage synthetic data generation to build OfficeQA Pro V2, allowing us to scale diverse, verified questions and answers. Combining these synthetic data techniques with our understanding of enterprise workflows enables us to rapidly build new benchmarks to make progress on the tasks our customers care about, like grounded reasoning. * An optimized agent harness can dramatically improve performance. Out-of-the-box agents averaged only 26.0% accuracy on OfficeQA Pro V2, while Databricks’ Genie delivered a 92% relative improvement on average across matched models and achieved up to 60% accuracy using the same models. Despite these gains, significant headroom remains in grounded reasoning.
