Skip to content
All topics

MLflow

Recent items mentioning MLflow across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.

60 recent items15 releases14 news18 videos13 community threads

What is MLflow?

MLflow is an open source platform for managing the lifecycle of machine learning models and, increasingly, LLM applications and agents. It tracks experiments (parameters, metrics, artifacts, code versions), packages models and agents in a standardized format, and registers them for versioned deployment. Databricks offers a fully managed and hosted version that builds on the open source experience, with the Model Registry integrated into Unity Catalog.

The problem it solves is losing track of what produced what. Without lifecycle tracking you can't reproduce a model, compare candidate versions, or explain an agent's behavior in production. MLflow Tracing records the inputs, outputs, and metadata of each intermediate step an application takes, which turns debugging and auditing from guesswork into reading a trace.

MLflow 3 moved the project's center of gravity toward generative AI: observability, evaluation with built-in LLM judges, and prompt management for agents and LLM apps, plus new concepts like Logged Models and Deployment Jobs. On Databricks, the managed version adds enterprise features such as encryption with customer-managed keys.

Is MLflow free?

The open source project is free to self-host. On Databricks you get a fully managed and hosted version that builds on the open source experience and adds enterprise security, disaster recovery, encryption with customer-managed keys, and tight Unity Catalog integration.

What's new in MLflow 3 compared to MLflow 2?

MLflow 3 centers on GenAI: end-to-end tracing for agent-based systems, evaluation with built-in LLM judges for safety, relevance, and correctness, and centralized prompt management with versioning and Unity Catalog integration. It also introduces the LoggedModel entity for organizing and comparing model variants, plus Deployment Jobs, with automatic tracing integrations for many popular GenAI libraries and frameworks.

Is MLflow only for classical machine learning?

No. MLflow Models is a standardized format for packaging both machine learning models and AI agents, and MLflow Tracing records the inputs, outputs, and metadata of each step a GenAI application takes. Databricks positions MLflow evaluation and monitoring for agents and ML models alike.

How does MLflow relate to Unity Catalog?

On Databricks, the MLflow Model Registry is integrated with Unity Catalog, so models are versioned and governed alongside your other data and AI assets. That gives you centralized governance and model lineage tracking, plus the ability to access models across workspaces and discover them for reuse.

Sources: MLflow on Databricks (Databricks docs) · MLflow 3 for GenAI (Databricks docs) · MLflow 3 release notes (mlflow.org)

What's happening in MLflowAI synthesis · updated 9h ago

MLflow 2.11.5 introduces an opt-in mechanism to route Unity Catalog model registry artifact uploads and downloads directly through the Databricks SDK Files API 10. Meanwhile, Databricks is standardizing MLflow execution tracing for AI agents managed under Unity AI Gateway 2, ahead of MLflow 3.17's planned integration with TypeSafe's Jev model for high-throughput QA evaluation 7.

Generated daily from the 10 most recent items mentioning MLflow. Click any [N] to jump to the source.

Reddit

How to fit ~1GB+ embedding model into a 2Gi Kubernetes pod? Getting OOMKilled

Hi, Deploying a FastAPI to Kubernetes that uses a multilingual sentence-transformers embedding model (ONNX backend, CPU only). My source data lives in a Delta table in Databricks, and the app reads from it and generates embeddings. The app takes user input (text) at embeds it, and compares it against stored embeddings for semantic similarity. So at least the query embedding has to happen live. Pod resources: - CPU: 1 request / 2 limit - Memory: 2Gi request = 2Gi limit (platform policy requires memory request and limit to be 1:1) Current docker setup: - Multi-stage Docker build (python:3.11-slim) - CPU-only PyTorch - The model is downloaded at build time, so it's baked into the image Image breakdown: - HF model cache: ~1.1 GB - torch: ~650 MB - pyarrow, scipy, transformers, pandas: ~100–150 MB each - Plus sklearn, onnxruntime, mlflow, and others The pod gets OOMKilled at startup or shortly after. With a ~1GB model, torch, pandas/pyarrow, and the ONNX runtime session all in one process, I think I'm just over 2Gi. What I'm trying to figure out on where do you store large models in production? Would love to hear what setups have worked for you. Thanks! submitted by /u/runningnozone [link] [comments]

00runningnozonetoday
Databricks CommunityMachine Learning

MLflow traces accepted by StartTraceV3 (200 OK) but never stored: GetTrace returns NOT_FOUND

002d ago
Reddit

Workshop on bringing systematic evaluation and MLflow tracking to LLM apps, Oct 3

If you're already using MLflow for experiment tracking on the ML side, there's a decent chance your LLM work hasn't caught up to the same standard yet. Most teams are still hand-tuning prompts and eyeballing outputs while everything else in the stack gets versioned and tracked properly. Serj Smorodinsky and Brett Kennedy, co-authors of a book on LLM applications, are running a live 3-hour workshop on Oct 3 that applies that same rigor to LLM development. What's covered: Programming LLM behavior with DSPy signatures and modules instead of hand-written prompt strings Building a baseline classifier live, from a real task Constructing an evaluation dataset with task-specific metrics so "better" is measurable, not a feeling Reading failure patterns directly out of the eval results Few-shot and instruction-level optimization applied systematically on top of the baseline Experiment tracking and LLM trace management through MLflow, so every run is reproducible and comparable Saving and reusing optimized DSPy programs across projects Communicating LLM reliability to stakeholders, which usually gets skipped entirely It's basically the MLflow discipline this community already applies to models, extended to prompts and LLM pipelines. Aimed at people already past "which model should I use" and into "how do I make this reliable." Full details here. submitted by /u/camerongreen95 [link] [comments]

00camerongreen954d ago
Reddit

Laya off the benchmark: can a zero-shot decision model route real SQL traffic?

Weekend-ish experiment on Databricks. One table, TPC-H orders, 15M rows, living in two places at once: Lakebase (Databricks' managed Postgres), a continuously synced copy with a btree index on the key. Sub-100ms point lookups, useless for a GROUP BY over 15M rows. Delta behind a serverless SQL Warehouse. Great at scans and aggregations, slow at fetching one row. The synced table is the nice part: native continuous Delta to Lakebase Postgres (needs a PK + CDF), so it's one dataset under one Unity Catalog, replication handled by the platform instead of a homemade pipeline. Then I put laya (convaiinnovations/laya, a non-autoregressive zero-shot decision model, vanilla, no fine-tuning) behind a FastAPI Databricks App. It reads each query and picks the engine. MLflow traces input to decision to execution. I submitted by /u/Limp-Park7849 [link] [comments]

00Limp-Park78491w ago
Databricks CommunityCommunity Articles

Building Custom Agents on Databricks:LangGraph, Atlan-Grounded Routing, and Per-Run Cost with MLflow

002w ago
Reddit

Could a “Data → Agent” composer be useful for Databricks?

I've been thinking about a gap between Databricks data and agent frameworks. Databricks already has a lot of the building blocks: - Unity Catalog - Genie / Genie Agents - MCP - Vector Search - AI Gateway - Agent skills/tools - MLflow - Omnigent / Kasal And tools like Omnigent and Kasal already solve a lot of the agent orchestration/execution side. But I'm wondering about the step before that: What if a customer already has a large, curated and governed Databricks data estate — how do we turn that data estate into an agent-ready configuration without manually wiring everything together? Something like: Existing Databricks Data Estate ↓ Data-to-Agent Composer ↓ ┌────────┼─────────┐ ↓ ↓ ↓ Domains Semantics Metrics ↓ ↓ ↓ Genie MCP Skills └────────┼─────────┘ ↓ Agent Configuration ↓ Omnigent / Kasal ↓ Agent The idea wouldn't be to build another chatbot or another agent framework. It would be a Databricks-native composition/bootstrapping layer that understands an existing Unity Catalog/data estate and generates the pieces needed for agents to work with that data — domain boundaries, semantic context, approved tools, Genie configuration, MCP exposure, skills, policies, evaluation setup, etc. In other words: Kasal/Omnigent: Agent → Tools/Data Proposed layer: Data Estate → Agent I'm curious if this is already solved somewhere in the Databricks ecosystem, or if people are currently doing this manually when building enterprise data agents. Would love to hear how others are approaching the “existing data estate → production-ready data agent” problem. submitted by /u/imsuryya [link] [comments]

00imsuryya2w ago
Databricks CommunityCommunity Articles

Building Custom Agents on Databricks:LangGraph, Atlan-Grounded Routing, and Per-Run Cost with MLflow

003w ago
Databricks CommunityMachine Learninganswered

Difference between Workspace and Unity Catalog experiments when using MLflow autologging?

001mo ago
Reddit

How are you actually setting up AI/LLM evals in Databricks end-to-end? Looking for a step-by-step production workflow

We have a product where Databricks is our backend , and we’re now trying to properly evaluate and improve the quality of the AI-generated answers in our application. I’m looking for advice from people who have actually implemented LLM/GenAI evaluations in Databricks in production . I’d really appreciate an end-to-end, step-by-step explanation of how you would set this up from scratch. Specifically, how would you approach: Define what a “good answer” means Accuracy / correctness Relevance Completeness Groundedness / faithfulness to our data Hallucination rate Citation/source correctness Following user instructions Response consistency Latency and cost Create a benchmark / golden dataset Should we manually create a set of representative user questions? How many questions are enough to start? Should each question have an expected answer? Should we store expected SQL/results, expected sources, or just an expected natural-language response? Where should this benchmark dataset live in Databricks? How do you keep it updated as the product evolves? Set up automated evaluations What Databricks/MLflow tools should we be using today? MLflow evaluation? LLM-as-a-judge? Custom scorers? Human evaluation? How do you combine these rather than relying on one score? Evaluate RAG / data-grounded answers Our AI answers questions based on enterprise data in Databricks. How do you separately evaluate: Retrieval quality Whether the correct tables/documents were selected Context relevance Groundedness Final answer correctness Whether the model invented something that wasn’t in the retrieved data Evaluate text-to-SQL / analytics questions If a user asks something like: “What was average building occupancy last month?” should we evaluate: Generated SQL Tables selected SQL execution result Final natural-language answer separately? What is the recommended architecture for this? Guardrails Where should guardrails sit in the architecture? For example: Prevent hallucinated numbers Prevent querying unauthorized tables Detect PII Prevent prompt injection Enforce tenant/user permissions Block unsupported questions Force answers to cite their source Return “I don’t know” when confidence is too low Should guardrails be part of evaluation, inference, or both? Production monitoring Once this is live, what should we log for every AI request? For example: User question → retrieved context → generated SQL/tool calls → query result → final response → model → prompt version → latency → tokens → cost → evaluation scores → user feedback Is that roughly the right model? Regression testing When we change: System prompt Model Retrieval strategy SQL generation logic Tools Temperature Data sources how do you automatically run the benchmark again and determine whether the new version is actually better? Do you set minimum score thresholds before allowing something to deploy? Human feedback How are people incorporating thumbs-up/down or analyst review into their evaluation datasets? Do production failures automatically become new benchmark cases? Making answers more precise This is ultimately my main goal. If our AI currently gives an answer that is “mostly correct,” what is the systematic process for figuring out why it isn’t fully correct? Is the best workflow something like: Production traces → identify failure → categorize failure → add to benchmark → improve retrieval/prompt/tool → run eval → compare against baseline → deploy → monitor Or is there a better approach? I’m especially interested in what the ideal Databricks-native architecture looks like: User Question → Agent / LLM → Retrieval / SQL / Tools → Databricks data → Response → MLflow tracing → Automated evaluators → Benchmark dataset → Regression testing → Production monitoring If you’ve implemented something like this, I’d love to know what you would build first, second, third, etc. Even a practical example like: Week 1: create benchmark Week 2: add tracing + scorers Week 3: add regression tests Week 4: add guardr […truncated]

00New_Championship39291mo ago
Databricks CommunityMachine Learning

E2E MLOps Part 1: How to build and govern models with AutoML, MLflow, and Unity Catalog

002mo ago
Databricks CommunityTechnical Blog

Multi-Agent Supervisor for Hybrid Retrieval with Agent Bricks and MLflow

004mo ago
RedditGeneral

Context Engineer Associate Beta Ex︁am + free attempt at DAIS

Context engineering is quickly becoming one of the key skills for building reliable AI agent systems. Databricks has just introduced the **Databricks Context Engineer Associate** **Ex︁am**, focused on designing, assembling, and governing the information AI agents receive at inference time - including prompts, retrieval systems, memory, tools, governance, and evaluation. The ex︁am is currently available as a **live beta at Data + AI Summit 2026**, and Databricks states that **one free onsite exam attempt will be offered during Summit**. Walk-ins only, one per attendee. Great opportunity for anyone working with GenAI, AI agents, Vector Search, Unity Catalog, MLflow, MCP, or Lakebase. [https://www.databricks.com/learn/certification/context-engineer-associate](https://www.databricks.com/learn/certification/context-engineer-associate)

76szymon_dybczak4mo ago
RedditDiscussion

Agent Deployment Pattern on Databricks App

Hi all, I'm figuring out what the standard pattern is to deploy an agent to Databricks App. Here's what I found and it is not that clear-cut: \- Pattern 1 (the conventional MLOps way): prepare agent script -> log agent (model-from-code) -> register agent -> deploy to Serving Endpoint \- Pattern 2 (the app template way): prepare agent script -> prepare agent server script (mlflow AgentServer) -> customize or use frontend codes -> deploy to Databricks App From Databricks documentation, pattern 2 is the recommended way to go for a multi-turn, long-running agent, and I can understand why. They even include a guide to migrate the agent from pattern 1 to pattern 2. However, it appears that we "break" the model life cycle process. While pattern 1 can retain all lineage from your experiment, run, catalog all the way to the endpoint, pattern 2 does not. The app becomes an isolated program. If I want to maintain the MLOps practice, I am stopped at the agent registration step, after which I have no pathway from the registered agent to the app. Am I understanding this correctly or there is something missing for pattern 2? Is there a native way to deploy a registered agent from unity catalog to an app? P.S. The training materials for the GenAI Engineering Associate cert is showing pattern 1.

44Careless_Handle81124mo ago

Get Tuesday's version of this

Tracking MLflow? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.

Read past issues first