Cost Optimization
Recent items mentioning Cost Optimization across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.
Enterprises are unlocking dramatic efficiency gains by modernizing compute models, exemplified by Octopus Energy cutting data pipeline costs 50x through Databricks Serverless and Delta Lake Change Data Feed 2, and NBCUniversal reducing spend by 30% by moving from slot reservations to dedicated job compute 1. For agent and inference deployments, practitioners comparing Databricks Apps and Model Serving can leverage Model Serving specifically when scale-to-zero capabilities and high-throughput governance are required to manage costs 3.
Generated daily from the 3 most recent items mentioning Cost Optimization. Click any [N] to jump to the source.
NBCUniversal’s Seamless Migration: Unlocking Scalable Analytics with Databricks
NBCUniversal cut costs by 30% by moving from a slot-based reservation model to Databricks' dedicated job compute, letting parallel pipelines scale independently and consistently hit SLA targets. Partner EXL executed the migration in phases with custom accelerators for code conversion, data migration, and automated validation, giving NBCUniversal a unified platform for ML development, real-time analytics, and collaborative data engineering.
The Open-Weight Revolution: A Game Changer for Our LLM Cost Optimization Odyssey
Snowflake to Databricks Migration in 12 weeks and cut cost per run by ~77%. AMA.
Lovelytics wrapped up a Snowflake-to-Databricks migration; 847 DBT models, 35 Info Mart tables, \~77% lower cost per run on a 2XL warehouse. **TL;DR What helped:** * Treated the migration as engineering, not translation. Each dbt model was tested in isolation, not just row counts vs Snowflake. * Routing macro to resolve cross-layer references at runtime, so the same codebase could read from Snowflake, federated Snowflake, and Unity Catalog without forking logic. * Dual model trees in one repo, which let the migration stay in lockstep with live Snowflake changes. * Script-generated wave selectors enabled parallel builds while preserving dependency order. * Used reference-slice validation subsets vs. waiting on full mart refreshes. **TL;DR Cost reduction:** * Reworked joins to use narrow staging dimensions instead of wide marts where possible. * Added incremental predicates to reduce MERGE target scans. * Split wide models into parallel sub-models where the dependency graph allowed it. * Copied static reference data into Delta instead of repeatedly reading it through federation. * Loaded static copies into Delta rather than reading via federation (predicate pushdown is poor). Happy to go into the gotchas: HASH() not being portable, Snowflake MERGE tolerating duplicate keys that Delta doesn't, NULL ordering, and timestamp handling. AMA [Full Blog Post](https://community.databricks.com/t5/technical-blog/partner-blog-847-models-12-weeks-77-less-inside-r1-s-snowflake/ba-p/157284)
Scaling for MHHS: how Octopus Energy achieved a 50x cost reduction in margin data engineering
Octopus Energy achieved a 50x cost reduction in their margin data engineering pipelines by re-architecting on Databricks for UK MHHS regulation. They leveraged Delta Lake Change Data Feed and Databricks Serverless to process 48x more data at a fraction of the original cost, improving freshness from weekly to daily.
[ebook] The Guide to Databricks Cost Optimization
[](https://www.reddit.com/r/databricks/?f=flair_name%3A%22General%22)Most cost optimization guides tell you to "right-size your clusters." Cool. Which ones? By how much? This guide actually answers that. Five core strategies, written by a Principal Data Engineer who's built production Databricks environments, for the people who own the bill. Download it for free (Unlike Your Databricks Bill.) [https://c.select.dev/guide-databricks-cost-optimization?utm\_source=reddit&utm\_medium=organic&utm\_campaign=databricks\_26](https://c.select.dev/guide-databricks-cost-optimization?utm_source=reddit&utm_medium=organic&utm_campaign=databricks_26)
[ebook] The No-BS Guide to Databricks Cost Optimization
Most cost optimization guides tell you to "right-size your clusters." Cool. Which ones? By how much? This guide actually answers that. Five core strategies, written by a Principal Data Engineer who's built production Databricks environments, for the people who own the bill. Download it for free (Unlike Your Databricks Bill.) [https://c.select.dev/no-bs-guide-to-databricks-cost-optimization?utm\_source=reddit&utm\_medium=organic&utm\_campaign=databricks\_26](https://c.select.dev/no-bs-guide-to-databricks-cost-optimization?utm_source=reddit&utm_medium=organic&utm_campaign=databricks_26) [](https://www.reddit.com/submit/?source_id=t3_1tdyfhd&composer_entry=crosspost_prompt)
Databricks cluster launch from external application
I want to launch a Databricks cluster from an external application. How can this be achieved, and what parameters need to be passed from the external application? Background: The user will already have the data ready for processing. Before execution, we want to provide multiple cluster configuration options based on the data volume. For example: For 50 GB → launch a smaller cluster For 100 GB → launch a medium cluster For 200 GB → launch a larger cluster ... Up to 1000 GB → launch a high-capacity cluster Based on the user’s selection, the external application should trigger Databricks to launch the appropriate cluster configuration and execute the workload. We would like to understand: The best approach to implement this architecture The required APIs/services to use The parameters that should be passed from the external application Recommended practices for dynamic cluster sizing, cost optimization, and workload execution in Databricks Any help/guidance. Thanks in advance.
NewsDatabricks Apps vs Model Serving: Authentication, Cost, and Performance Compared
Databricks Apps are now the recommended first choice for deploying agents due to their flexibility in handling full-stack applications with multiple components, offering faster iteration and local testing compared to Model Serving. Model Serving remains suitable for use cases prioritizing high QPS, governance features like AI Gateway, inference tables, and guardrails, or when scaling to zero is acceptable for cost optimization.
NewsGPU Accelerated Spark Connect
This video demonstrates how to accelerate Spark Connect using GPUs for both Spark SQL and ML workloads. It details the architecture, deployment, and benchmark results showing significant speedups and cost savings compared to CPU-only execution.
NewsSaving Millions From Millions: Navigating Towards Cost-Efficiency in Pinterest's Spark Jobs
Pinterest built a three-layer cost optimization system for Spark jobs using observability platforms (collecting real-time metrics and event logs), platform innovations (Apache Celeborn remote shuffle service and Kubernetes burst-aware memory allocation), and fine-grained management tools. Production deployments achieved 30-50% cost reductions per job through resource tuning while improving cluster utilization.
NewsInnovating Retail Data: Unilever’s Transformation with Databricks DLT
Unilever replaced its legacy data pipelines with Databricks Delta Live Tables using a medallion architecture, enabling serverless streaming, automated data quality checks, and unified governance through Unity Catalog. The migration delivered 25% infrastructure cost reduction, 200-500% faster data processing, and real-time analytics at scale.
NewsFrom Code to Insights: Leveraging Advanced Infrastructure and AI Capabilities
Insulet migrated its data infrastructure from Azure to Databricks, implementing a data mesh architecture with medallion layers, Unity Catalog governance, and Delta Sharing for partner data access. The company achieved real-time analytics, GXP-compliant deployments, and FinOps cost optimization, demonstrating how to balance innovation with operational efficiency.
NewsSponsored: Kyvos | Analytics 100x Faster Lowest Cost w/ Kyvos & Databricks, Even on Trillions Rows
Get Tuesday's version of this
Tracking Cost Optimization? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.






