본문으로 건너뛰기
← 전체 뉴스
Databricks Blog2026년 8월 18일

Evaluating AI Agents Live at the Grounded Reasoning Cup

요약

Stanford's team won the Grounded Reasoning Cup with 63.3% accuracy on OfficeQA Pro V2, a new 120,000-page U.S. Treasury document benchmark, using an end-to-end agent optimization approach combining reusable skills, document-representation fallbacks, and adaptive verification. Across all 11 academic teams, out-of-the-box frontier agents averaged under 30% accuracy, showing that agent performance tuned on one benchmark doesn't reliably generalize to a new corpus.

* The Grounded Reasoning Cup challenged 11 academic teams to apply agents developed on OfficeQA Pro to OfficeQA Pro V2, a newly released benchmark built from approximately 120,000 pages of U.S. Treasury documents. * Results showed that generalization cannot be assumed. Approaches developed on a familiar benchmark did not always transfer reliably to a new corpus, and out-of-the-box frontier agents averaged less than 30% accuracy. * Stanford’s winning team achieved 63.3% accuracy through an end-to-end agent optimization strategy that combined a library of reusable skills, targeted document-representation fallbacks, and adaptive verification.

관련 기사

News

Databricks Feature Store가 1초 미만의 신선도로 피처를 제공하는 방법

databricks-blog14h ago
News

'프로토타이핑 세금'이 AI 로드맵을 무너뜨리고 있다

databricks-blog17h ago
News

data warehouse에서 AI_Functions 활용하기: 주요 사용 사례

databricks-blog3d ago
News

Scottish Water가 Databricks Genie로 설비투자 데이터를 대화형으로 만든 방법

databricks-blog4d ago