Kubernetes
Recent items mentioning Kubernetes across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.
Databricks' AI SRE now triages over 2,000 daily investigations across 1,500 Kubernetes clusters spanning 70+ regions and three clouds 3. Meanwhile, practitioners are actively migrating workloads from Databricks Jobs to Kubernetes-based Spark jobs 1.
Written September 2026 from the 3 most recent items mentioning Kubernetes at that time. It refreshes when this topic next has enough new material. Click any [N] to jump to the source.
Databricks Jobs to Kubernetes Spark Jobs migration
I'm discussing some architectural changes with my data team, and one of them is about cutting costs. The idea is to migrate Databricks Jobs to Spark jobs on Kubernetes. We've done a PoC and it works well. But despite the clear cost savings, I'm getting pushback from management and they're resisting the change. Can someone explain the reason behind this? It seems clear that you pay double for running jobs in Databricks, and you could save a lot of money. Yes, doing this means we bypass Unity Catalog, but we could convert only the Bronze and Silver layers. submitted by /u/worldpwn [link] [comments]
MLflow 3.16.0
MLflow 3.16.0 makes the redesigned trace explorer the default interface, introducing natural-language custom trace views via the MLflow Assistant, session grouping for multi-turn conversations, and span links. The update also adds Unity Catalog model service support for built-in evaluation judges, per-user AI Gateway budget policies, and fail-closed authorization by default.
Data Engineers, what does your actual day-to-day work look like? And what should I learn next?
I’m currently trying to transition deeper into Data Engineering and would really appreciate some perspective from people who are already working in the field. I have 1.3 yrs experience as a Junior Python Developer. What I want to do is slowly transform into a Data Engineer. How would you suggest my choice? Basically what I do is make web scraping scripts to get the data from web and give the data in excels. Our company is currently not using git or CI/CD or anything like that. The problem I’m running into is that when I look at Data Engineering jobs on Naukri, LinkedIn, etc., the requirements seem endless. One job asks for Python, SQL, Airflow and AWS; another wants Spark, Kafka and Databricks; another wants Snowflake, dbt, Terraform, Kubernetes, CI/CD, etc. It becomes difficult to understand what I should actually prioritize. So I’d like to hear from people who are actually working as Data Engineers . What does your day-to-day work look like? What kind of problems do you solve, what technologies do you use regularly, and which skills have turned out to be genuinely important in your job? More importantly, based on my current experience, what would you suggest I improve or learn next to become a stronger candidate for Data Engineering roles? Are there any gaps that you think I should focus on, or technologies/concepts that are worth learning through projects rather than just studying theoretically? I’m not really looking for a generic “learn SQL → Python → Spark → AWS” roadmap. I’m more interested in understanding the reality of the job and getting advice from people who have actually gone through the transition. If you’re a Data Engineer with 1–5+ years of experience , I’d especially appreciate your perspective. Even a short description of what you work on and what you wish you had learned earlier would be extremely helpful. Thanks in advance! submitted by /u/Stunning-Space8032 [link] [comments]
How Databricks Uses AI to Accelerate Incident Investigation
Databricks' AI SRE now handles over 2,000 daily investigations across 150+ teams, helping engineers troubleshoot 100s of microservices spread across 1,500 Kubernetes clusters in 70+ regions and three clouds. Rather than relying on black-box reasoning, the system uses composable agentic runbooks and ties every diagnostic recommendation to verifiable raw evidence, embedding a context-first approach into its design.
MLflow 3.15.0 introduces an MCP Registry for registering and sharing Model Context Protocol servers, enhances the Assistant with multi-provider LLM support and per-session token usage tracking, and enables proxy-less artifact transfers via presigned URLs to reduce server load and timeouts on large files. Additional improvements include sharable Runs table views, multi-modal image attachments for LLM judges to evaluate vision tasks, and numerous bug fixes across tracing, evaluation, gateway, and UI components.
MLflow 3.13.0 introduces Role-Based Access Control with Admin UI, automatic trace archival to S3, and one-click observability for Claude Code and other coding agents. Breaking changes include a redesigned permission system (legacy APIs removed), MLServer removal from pyfunc serving, and requirement for MLFLOW_ALLOW_FILE_STORE=true flag for local file-based stores.
MLflow 3.13.0rc0 completely overhauls Role-Based Access Control with unified permission APIs and a new Admin UI, and integrates Claude Code, OpenAI, Ollama, and OpenClaw as native assistant providers in the AI Gateway. The release adds trace archival with seamless retrieval, GenAI agent stress-testing, Kubernetes Helm chart support, and database replica routing for horizontal scaling.
MLflow 3.11.1 introduces AI-powered issue detection in traces, AI Gateway budget alerts and spending controls, trace graph visualization, native Databricks gateway provider, and pickle-free model serialization. TypeScript SDK packages are now @mlflow-scoped and LiteLLM is no longer required for GenAI evaluation.
Enterprise-Scale MLflow Operations and Security Practices at LY Corporation
LY Corporation deployed a managed MLflow architecture on Kubernetes handling over 600,000 daily requests across 40 services while maintaining strict data isolation. By leveraging Athenz sidecar proxies and SPIFFE-based mTLS for automated OAuth 2.0 authentication, the platform secures machine-to-machine communication from training pods without modifying open-source MLflow.
UnityCatalog 0.3.0
UnityCatalog 0.3.0 adds support for Spark 4.0 and Delta 4.0 with new credentials and external locations APIs enabling flexible external storage management. The release includes Kubernetes deployment via Helm charts and fixes for Delta time-travel SQL queries.
NewsSaving Millions From Millions: Navigating Towards Cost-Efficiency in Pinterest's Spark Jobs
Pinterest built a three-layer cost optimization system for Spark jobs using observability platforms (collecting real-time metrics and event logs), platform innovations (Apache Celeborn remote shuffle service and Kubernetes burst-aware memory allocation), and fine-grained management tools. Production deployments achieved 30-50% cost reductions per job through resource tuning while improving cluster utilization.
NewsInternet-Scale Analytics: Migrating a Mission Critical Product to the Cloud
Akamai migrated its high-volume internet security analytics platform from an on-premises Hadoop architecture to a cloud-native setup using Azure blob storage, Kafka, Kubernetes, and Databricks. The resulting architecture decouples raw data storage from streaming metadata via Kafka, utilizing custom shared libraries and sharding strategies to achieve low-latency ingestion, cost-effective scaling, and improved performance.
Get Tuesday's version of this
Tracking Kubernetes? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.




