Lakehouse Federation
Recent items mentioning Lakehouse Federation across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.
Organizations are deploying Lakehouse Federation alongside Unity Catalog as a governed context layer for AI agents and natural language querying across disparate systems, eliminating data migration in enterprise R&D and telecom workflows 234. In the field, practitioners querying federated Snowflake sources are encountering operational hurdles, specifically with large result sets failing to download from internal stages 1.
Written September 2026 from the 4 most recent items mentioning Lakehouse Federation at that time. It refreshes when this topic next has enough new material. Click any [N] to jump to the source.
Lakehouse Federation (Snowflake) — large query results fail to download from internal stage
Backstage with Lakebase, part 3
Databricks Lakehouse Federation enables teams to query live Backstage operational data directly alongside Unity Catalog billing tables with zero ETL data movement. Learn how to configure the required Postgres authentication workaround for Lakebase and leverage isolated compute to monitor the exact DBU cost of ephemeral development branches.
Why R&D Data Belongs in the Lakehouse - and Why Agents Need It There
Cellcentric built an industrial Data Hub on Databricks using Unity Catalog and Lakehouse Federation, establishing a governed context layer that unifies scattered R&D data for employees and AI agents via an MCP server. By treating documentation as a first-class quality metric, this architecture delivers AI-ready data products that accelerate complex R&D investigations from weeks to days.
How to Deploy Lakehouse Federation Using DABs from the Dev Environment to the Stage Environment
Talk to all your data, wherever it lives
Lakehouse Federation is now available, allowing you to query data across all sources without migration delays. Unity Catalog serves as the single source of truth for both federated and managed data, enabling secure AI workloads and natural language querying.
When Lakehouse Federation "loses" your tables: the silent space-in-name trap
AI readiness in telecommunications
Telco AI initiatives stall at production scale due to data debt, not model quality; Databricks Unity Catalog provides the semantic layer and governance needed to bridge this gap. It unifies disparate systems via Lakehouse Federation, offering AI agents rich context and enabling end-to-end governance for regulatory compliance and accurate operational tasks.
Tutorial: joining Lakebase OLTP and Unity Catalog Delta in one federated SQL query, with a text-to-SQL agent
I wanted to answer one natural-language question that crosses two stores: "Show me orders from high-value customers that shipped last week." The order status is live in Postgres (OLTP). The customer-segment definition lives in a Delta gold table (OLAP). For years this was a two-system query — pull from one, dump to CSV, join in pandas, regret your life. With Lakebase (managed Postgres) sitting in the same Databricks workspace as Unity Catalog, Lakehouse Federation makes that a single SQL statement. And once it's a single statement, you can put a text-to-SQL agent on top of it and stop writing the query by hand. This post walks the whole thing end-to-end: data setup, federation wiring, the agent pipeline, the safety guardrail, and three swappable LLM backends (Gemini via Google AI Studio, GPT-4.1 / o4-mini via GitHub Models, and self-hosted vLLM on a Shadeform H100). The data setup I used the Olist Brazilian e-commerce dataset (Kaggle, \~100MB). Loaded twice on purpose: data/raw/\*.csv (customers, orders, items, products, payments) │ ├─► load\_oltp.py ─────► Lakebase Postgres │ (5 normalized tables — the live source-of-truth) │ └─► data/processed/\*.csv (cleaned, dedup'd) │ └─► UC Volume → build\_olap.py ─► Unity Catalog Delta (gold aggregates: customer\_segments, category\_performance, revenue\_aggregates) The OLTP side is what your app would write to in production. The OLAP side is what your analytics jobs would build nightly. We want to query both from the same agent without it knowing or caring which is which. Wiring Lakehouse Federation Three steps from a clean workspace. * Create the connection in Unity Catalog. `CREATE CONNECTION lakebase_olist` `TYPE postgresql` `OPTIONS (` `host 'YOUR-LAKEBASE-INSTANCE.database.cloud.databricks.com',` `port '5432',` `user 'chandank@becloudready.com',` `password secret('lakebase', 'pat-token'),` `trustServerCertificate 'true'` `);` * Create the foreign catalog. This is where the database option lives — not on the connection. (This trips up most people first time, including me.)`CREATE FOREIGN CATALOG lakebase_olistUSING CONNECTION lakebase_olistOPTIONS (database 'olist');` I wanted to answer one natural-language question that crosses two stores: The order status is live in Postgres (OLTP). The customer-segment definition lives in a Delta gold table (OLAP). For years this was a two-system query — pull from one, dump to CSV, join in pandas, regret your life. With Lakebase (managed Postgres) sitting in the same Databricks workspace as Unity Catalog, Lakehouse Federation makes that a single SQL statement. And once it’s a single statement, you can put a text-to-SQL agent on top of it and stop writing the query by hand. This post walks the whole thing end-to-end: data setup, federation wiring, the agent pipeline, the safety guardrail, and three swappable LLM backends (Gemini via Google AI Studio, GPT-4.1 / o4-mini via GitHub Models, and self-hosted vLLM on a Shadeform H100). # The data setup I used the Olist Brazilian e-commerce dataset (Kaggle, \~100MB). Loaded twice on purpose: data/raw/*.csv (customers, orders, items, products, payments) │ ├─► load_oltp.py ─────► Lakebase Postgres │ (5 normalized tables — the live source-of-truth) │ └─► data/processed/*.csv (cleaned, dedup'd) │ └─► UC Volume → build_olap.py ─► Unity Catalog Delta (gold aggregates: customer_segments, category_performance, revenue_aggregates) The OLTP side is what your app would write to in production. The OLAP side is what your analytics jobs would build nightly. We want to query both from the same agent without it knowing or caring which is which. # Wiring […truncated]
How would you onboard legacy data stores that don't use OAuth into Databricks Unity Catalog?
Lakehouse federation typically uses OAuth, this the above qn.
Tutorials43 Lakehouse Federation or Query Federation in Databricks | Query External Database |Foreign Catalog
Lakehouse Federation in Databricks allows querying external databases like PostgreSQL directly without importing data, using connections and foreign catalogs with Unity Catalog governance and lineage support. The demo shows setting up a PostgreSQL connection, creating a foreign catalog, and executing federated SQL queries that join external tables.
EventsDatabricks Product Announcements at Data + AI Summit 2024
Databricks launched Delta 4.0 with lakehouse federation and open-source Unity Catalog, plus Databricks AI BI for conversational business intelligence powered by a learning system called Genie. The company also expanded Mosaic AI with serverless GPUs, zero-code model fine-tuning, an agent framework, improved ML Flow tracing, and the Mosaic AI Gateway for governance and auditability.
ReleasesAnnouncing Databricks Clean Rooms with Live Demo. Presented by Matei Zaharia and Darshana Sivakumar
Databricks Clean Rooms enables two parties to securely collaborate on private data, unstructured assets, and machine learning models without exposing sensitive information. The platform utilizes Delta sharing and Lakehouse Federation to allow cross cloud and cross platform joint analysis using Python, R, or SQL.
EventsData + AI Summit 2024 - Keynote Day 2 - Full
The Databricks Data plus AI Summit 2024 keynote announces the general availability of Delta Uniform and Lakehouse Federation, alongside the open-sourcing of Unity Catalog via the Linux Foundation. The presentations also demo duckDB integration with Delta Lake, introduce DuckDB version 1.0, and highlight new features in Delta 4.0 such as liquid clustering and the open variant data type.
Get Tuesday's version of this
Tracking Lakehouse Federation? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.

