Scaling and Operating a Large dbt Project on Databricks: IFCO's Data Team on Performance, Visibility, and Debugging
Summary
Tuning dbt incremental models with liquid clustering, dynamic file pruning, and deliberate merge strategies cut IFCO's core job runtime by over 60% and retired their nightly full refresh. By diagnosing real executed query plans and orchestrating models as discrete Databricks Jobs tasks via the open-source databricks-dbt-factory, the team gained per-model visibility, targeted reruns, and enforced testing.
Summary generated by brickster.ai. For the full article, follow the source link above.
More from Databricks Blog
Meta’s ads MCP server comes to Databricks: Put your customer intelligence to work in advertising campaigns
Meta’s ads MCP server is now available in the Databricks Marketplace, allowing marketers to use natural language in Genie to evaluate campaign performance alongside governed business data. Administrators can control access to the connection through Unity Catalog, with Unity Gateway governing tool calls and recording audit logs.
NEAREST BY Join: Scaling Vector Search in Databricks Runtime
Databricks Runtime now features NEAREST BY, a new SQL join that executes exact or approximate batch vector searches directly against your Lakehouse data. Powered by a fused Photon operator with a custom blocked GEMM kernel, it turns standard liquid-clustered Delta tables into partition-pruned vector indexes with no external vector store to sync or manage.
Unlocking Data Portability: Preventing Catalog Lock-in with REGISTER and UNREGISTER APIs
The new REGISTER and UNREGISTER APIs provide a standard method to safely transfer open table format management between catalogs without copying data stored in customer-owned buckets. By requiring the original catalog to explicitly relinquish control before a transfer, this workflow prevents catalog lock-in and eliminates dangerous split-brain scenarios.
How Genie Ontology powers product development at Databricks
Databricks relies on Genie Ontology and Genie One to drive product development, combining human-curated semantics and learned knowledge across first- and third-party sources to analyze adoption and forecast trends. The platform extends these internal workflows with capabilities including custom skills, scheduled tasks, shareable agents, and MCP writes to external tools.
In Capital Markets, the Buy Side Runs on NAV. Finance Protects the Fee.
Databricks Genie One introduces an AI coworker for asset management finance that leverages backend agents to automate NAV reconciliation and fund data ingestion across a governed business ontology. Designed to protect net fee margins under strict regulatory standards like Form PF and GIPS, it provides finance teams with real-time, traceable morning NAV calculations and audit-ready proof for every valuation.
How to choose your first Genie Agents for maximum impact
Rank candidate Genie Agent workflows in minutes using a five-criteria rubric across impact, demand, data readiness, scope, and governance to ensure adoption scales rather than stalls. Field examples demonstrate how to apply this framework in practice to classify use cases into build now, shape it first, and hold off categories.
