Skip to content
All news
PlatformDatabricks Blog·August 18, 2026·Databricks AI Research Team

Evaluating AI Agents Live at the Grounded Reasoning Cup

Summary

Stanford's team won the Grounded Reasoning Cup with 63.3% accuracy on OfficeQA Pro V2, a new 120,000-page U.S. Treasury document benchmark, using an end-to-end agent optimization approach combining reusable skills, document-representation fallbacks, and adaptive verification. Across all 11 academic teams, out-of-the-box frontier agents averaged under 30% accuracy, showing that agent performance tuned on one benchmark doesn't reliably generalize to a new corpus.

Summary generated by brickster.ai. For the full article, follow the source link above.

More from Databricks Blog