Skip to content
All topics

Lakehouse Federation

Recent items mentioning Lakehouse Federation across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.

14 recent items4 news5 videos5 community threads
What's happening in Lakehouse FederationAI synthesis · updated 28d ago

Organizations are deploying Lakehouse Federation alongside Unity Catalog as a governed context layer for AI agents and natural language querying across disparate systems, eliminating data migration in enterprise R&D and telecom workflows 234. In the field, practitioners querying federated Snowflake sources are encountering operational hurdles, specifically with large result sets failing to download from internal stages 1.

Written September 2026 from the 4 most recent items mentioning Lakehouse Federation at that time. It refreshes when this topic next has enough new material. Click any [N] to jump to the source.

Databricks CommunityData Engineering

Lakehouse Federation (Snowflake) — large query results fail to download from internal stage

001mo ago
Databricks CommunityData Engineering

How to Deploy Lakehouse Federation Using DABs from the Dev Environment to the Stage Environment

002mo ago
Databricks CommunityData Governance

When Lakehouse Federation "loses" your tables: the silent space-in-name trap

003mo ago
RedditTutorial

Tutorial: joining Lakebase OLTP and Unity Catalog Delta in one federated SQL query, with a text-to-SQL agent

I wanted to answer one natural-language question that crosses two stores: "Show me orders from high-value customers that shipped last week." The order status is live in Postgres (OLTP). The customer-segment definition lives in a Delta gold table (OLAP). For years this was a two-system query — pull from one, dump to CSV, join in pandas, regret your life. With Lakebase (managed Postgres) sitting in the same Databricks workspace as Unity Catalog, Lakehouse Federation makes that a single SQL statement. And once it's a single statement, you can put a text-to-SQL agent on top of it and stop writing the query by hand. This post walks the whole thing end-to-end: data setup, federation wiring, the agent pipeline, the safety guardrail, and three swappable LLM backends (Gemini via Google AI Studio, GPT-4.1 / o4-mini via GitHub Models, and self-hosted vLLM on a Shadeform H100). The data setup I used the Olist Brazilian e-commerce dataset (Kaggle, \~100MB). Loaded twice on purpose: data/raw/\*.csv (customers, orders, items, products, payments) │ ├─► load\_oltp.py ─────► Lakebase Postgres │ (5 normalized tables — the live source-of-truth) │ └─► data/processed/\*.csv (cleaned, dedup'd) │ └─► UC Volume → build\_olap.py ─► Unity Catalog Delta (gold aggregates: customer\_segments, category\_performance, revenue\_aggregates) The OLTP side is what your app would write to in production. The OLAP side is what your analytics jobs would build nightly. We want to query both from the same agent without it knowing or caring which is which. Wiring Lakehouse Federation Three steps from a clean workspace. * Create the connection in Unity Catalog. `CREATE CONNECTION lakebase_olist` `TYPE postgresql` `OPTIONS (` `host 'YOUR-LAKEBASE-INSTANCE.database.cloud.databricks.com',` `port '5432',` `user 'chandank@becloudready.com',` `password secret('lakebase', 'pat-token'),` `trustServerCertificate 'true'` `);` * Create the foreign catalog. This is where the database option lives — not on the connection. (This trips up most people first time, including me.)`CREATE FOREIGN CATALOG lakebase_olistUSING CONNECTION lakebase_olistOPTIONS (database 'olist');` I wanted to answer one natural-language question that crosses two stores: The order status is live in Postgres (OLTP). The customer-segment definition lives in a Delta gold table (OLAP). For years this was a two-system query — pull from one, dump to CSV, join in pandas, regret your life. With Lakebase (managed Postgres) sitting in the same Databricks workspace as Unity Catalog, Lakehouse Federation makes that a single SQL statement. And once it’s a single statement, you can put a text-to-SQL agent on top of it and stop writing the query by hand. This post walks the whole thing end-to-end: data setup, federation wiring, the agent pipeline, the safety guardrail, and three swappable LLM backends (Gemini via Google AI Studio, GPT-4.1 / o4-mini via GitHub Models, and self-hosted vLLM on a Shadeform H100). # The data setup I used the Olist Brazilian e-commerce dataset (Kaggle, \~100MB). Loaded twice on purpose: data/raw/*.csv (customers, orders, items, products, payments) │ ├─► load_oltp.py ─────► Lakebase Postgres │ (5 normalized tables — the live source-of-truth) │ └─► data/processed/*.csv (cleaned, dedup'd) │ └─► UC Volume → build_olap.py ─► Unity Catalog Delta (gold aggregates: customer_segments, category_performance, revenue_aggregates) The OLTP side is what your app would write to in production. The OLAP side is what your analytics jobs would build nightly. We want to query both from the same agent without it knowing or caring which is which. # Wiring […truncated]

42kchandank4mo ago
RedditDiscussion

How would you onboard legacy data stores that don't use OAuth into Databricks Unity Catalog?

Lakehouse federation typically uses OAuth, this the above qn.

11RazzmatazzLiving13234mo ago

Get Tuesday's version of this

Tracking Lakehouse Federation? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.

Read past issues first