MLflow 3.15 introduced immutable evaluation datasets, ensemble scoring, and serverless runtime fixes.
In August 2026, MLflow received two patch updates focused on evaluation workflows and runtime reliability. MLflow 3.15.1 corrected version parsing issues on Databricks Serverless and adjusted packaging behavior on ARM client images [12]. MLflow 3.15.2 introduced immutable evaluation dataset versions and the scorer_ensemble primitive for aggregating judge outputs [2]. The patch also resolved a Databricks telemetry deadlock and fixed edge cases in aligned judge execution and run status constraints [2].
Databricks tooling expanded its integration with MLflow storage and tracking. Both the Databricks SDK and platform job definitions gained support for configuring custom MLflow artifact locations within automated tasks [7, 8]. In addition, the Databricks AI Runtime command-line tool surfaced native MLflow experiment tracking alongside hardware utilization metrics for distributed LLM training runs across serverless GPU clusters [4].
AI agent observability and evaluation remained prominent themes throughout the month. Educational materials detailed how to capture agent execution steps with MLflow traces when orchestrating multi-tool systems [9], paired with guidance on evaluating agent skills [13]. Meanwhile, community discussions reflected practical demands from teams seeking established step-by-step blueprints to operationalize language model evaluations in production Databricks environments [1].
Everything cited
- [1]How are you actually setting up AI/LLM evals in Databricks end-to-end? Looking for a step-by-step production workflow community · 2026-08-30
- [2]v3.15.2 release · 2026-08-26
- [3]Building Governed Agents with Databricks video · 2026-08-19
- [4]AI Runtime CLI | Serverless GPU LLM Training video · 2026-08-18
- [5]How AI and Data Keep 2.3 Million Lawns Healthy | TruGreen & Databricks video · 2026-08-16
- [6]Databricks News: ZeroOps, DABs, Indexes, Genie, sandboxes, migration from PowerBI, secrets video · 2026-08-13
- [7]v0.145.0 release · 2026-08-12
- [8]v0.127.0 release · 2026-08-12
- [9]Building Agents on Databricks with Custom Apps and Omnigent video · 2026-08-07
- [10]Building Agents on Databricks with Custom Apps and Omnigent video · 2026-08-06
- [11]AI for Mental Health: Crisis Text Line and Databricks video · 2026-08-04
- [12]MLflow 3.15.1 release · 2026-08-03
- [13]Evaluating and Improving Agent Skills with MLflow news · 2026-08-02
A frozen monthly snapshot, generated from the brickster.ai archive and never rewritten. For the live view of this topic, see the MLflow hub. brickster.ai is an independent community project, not affiliated with Databricks, Inc.
