Scaling and Operating a Large dbt Project on Databricks: IFCO's Data Team on Performance, Visibility, and Debugging
Tuning dbt incremental models with liquid clustering, dynamic file pruning, and deliberate merge strategies cut IFCO's core job runtime by over 60% and retired their nightly full refresh. By diagnosing real executed query plans and orchestrating models as discrete Databricks Jobs tasks via the open-source databricks-dbt-factory, the team gained per-model visibility, targeted reruns, and enforced testing.
* Tune dbt incremental models so each run writes only the rows that changed, using liquid clustering, a deliberate merge strategy, and dynamic file pruning. That cut IFCO's core job runtime by over 60% and retired the nightly full refresh. * Diagnose slow models from the real executed query plan, not the compiled SQL, since that is where the true causes appear. IFCO packaged that diagnosis into a repeatable skill that runs on every model. * Run the dbt project as one task per model on Databricks Jobs with the open-source databricks-dbt-factory, not as one opaque job. That gives per-model visibility, targeted reruns, and enforced testing.
