Skip to content
All topics

Lakebase

Recent items mentioning Lakebase across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.

60 recent items4 releases17 news7 videos32 community threads

What is Lakebase?

Lakebase Postgres is a fully managed Postgres database built into the Databricks platform. It's real Postgres, so standard drivers and clients work, and instead of a fixed server you provision a project with branches and databases inside it.

Analytics tables in the lakehouse are a poor fit for application backends that need low latency reads and writes, and that's the gap Lakebase fills without leaving the platform. You can sync Unity Catalog tables into Lakebase so applications query them quickly, and store Postgres changes back as Delta tables for pipelines and audit (that direction is in Public Preview). Autoscaling, scale-to-zero, and instant branching keep development and test environments cheap to create.

Lakebase reached general availability on January 22, 2026, with automated backups and point-in-time recovery; read replicas and high availability with automatic failover are also available, the latter added in March 2026. As of September 2026 it tracks current Postgres releases, with major versions 16, 17, and 18 supported.

Is Lakebase generally available?

Yes. Lakebase reached GA on January 22, 2026, including autoscaling, scale-to-zero, instant branching, automated backups, and point-in-time recovery. One notable exception is storing Postgres changes as Delta tables with full change history, which is in Public Preview.

Is Lakebase real Postgres or just Postgres compatible?

It's actual managed Postgres, and standard Postgres drivers and clients work against it. As of September 2026 it supports Postgres major versions 16, 17, and 18.

How much does Lakebase cost?

Billing is usage based and began in January 2026, with snapshot storage billed separately since June 1, 2026. Scale-to-zero and autoscaling are the main levers for keeping idle branches and quiet databases cheap; Databricks publishes current rates per cloud and region on its pricing page.

Do I need Lakebase if I already have the lakehouse?

The lakehouse handles analytics, while Lakebase is for transactional application workloads that need low latency reads and writes. The two connect: you can sync Unity Catalog tables into Lakebase for fast serving, and capture Postgres changes back as Delta tables for downstream pipelines and audit.

Sources: Lakebase Postgres (Databricks docs) · Lakebase release notes (Databricks docs)

What's happening in LakebaseAI synthesis · updated 10h ago

Lakebase Postgres introduced metadata-driven, branch-based restores against decoupled immutable storage, slashing recovery times for 100 TB databases down to seconds 2. Supported by new Databricks CLI expiration flags 8, these isolated database branches now enable automated testing and preview environments for parallel AI coding agents 3, while Lakebase expands into real-time architectures as an online feature serving layer 1.

Generated daily from the 10 most recent items mentioning Lakebase. Click any [N] to jump to the source.

Reddit

Databricks Lakebase Search - GA Announcements

submitted by /u/Wise_Ear_4064 [link] [comments]

00Wise_Ear_40642d ago
Databricks CommunityData Engineering

Built my first Databricks App with Lakebase

002d ago
Reddit

Databricks Micro Apps, App Spaces and Genie App Generator

Databricks launched App Spaces and Serverless Micro Apps at the end of last week (in Beta). Over the weekend, I tested migrating two of my production Databricks Apps to App Spaces. App Migration 1: Blocked by Zero Egress My Data Portfolio Project Creator required internet access to pull data stacks from live job postings. Then, the LLM needs internet access to research for open data sources to use. In standard Databricks Apps, this runs cleanly. In App Spaces, there is zero external internet egress including for LLMs. The app can reach internal workspace resources, but it cannot touch the outside web. If your app relies on third-party APIs, external databases, or web scraping, App Spaces is a non-starter until Databricks opens network egress. Workload 2: Success on Internal FinOps My client-facing DBU cost observability app reads workspace usage data and writes directly to Lakebase. Because it requires zero external network calls, the migration worked. Both the micro app compute and Lakebase scale to zero when idle. Cold starts take roughly 30 seconds (though in beta, you occasionally need a quick browser refresh once it spins up). For internal, low-frequency administrative tools, this turns a continuous monthly compute bill into pennies. (+ App Spaces, Genie App Generator, and Micro Apps are free while in beta) The next part is less about App Spaces and more just general best practice for Databricks Apps that I see people miss. Stop Using Delta Lake as an OLTP Database Databricks Apps are software applications, not batch analytics notebooks. If your app writes application state, session data, or row-level CRUD directly into analytical Delta tables, you need to rethink that design. For app transactions, use Lakebase (serverless Postgres which also scales to 0). Both are governed under UC, but Lakebase gives your app the low-latency transactional engine that application engineering actually requires. My Verdict on the Beta App Spaces solves the idle compute problem that has plagued Databricks Apps since launch. But until Databricks allows us to deploy directly from existing Git repos and opens external network egress, it remains limited to internal-only use cases. Curious on other peoples experience... How has it been for others? submitted by /u/OkImprovement7010 [link] [comments]

00OkImprovement70104d ago
Reddit

Operational Responsibilities in DBX

Has anyone come across job responsibilities that aren't software developers (engineers, analysts, or scientists)? In other words, I'm wondering if there is such a thing that corresponds to a classic DBA position, but for this Databricks SaaS? Now that UC is here with managed catalogs, the technology is looking more like a conventional database platform. In fact the new lakebase is actually a conventional Oltp based on postgres (or LTAP if you prefer the latest marketing). It seems like the ecosystem is complex enough now to support some operational job positions, especially if the data estate is large. Anyone have a position that is more than 50 pct operational?) What about 100 pct operational, like a devops or DBA that focuses on databricks submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36971w ago
Reddit

Lakebase handles failover within a region, but cross-region disaster recovery is still in Private Preview. How are people covering that?

I was reading through Databricks' breakdown of what Lakebase takes over on the ops side. Patching, scaling, in-region failover, and point-in-time restore with 2 to 30 days of history all run automatically, while cross-region disaster recovery is Private Preview on AWS only, with manual failover and recovery procedures the customer owns. For anyone running Lakebase behind an app or for agent state, how are you handling a region outage today? I can see people copying scheduled snapshots elsewhere, replicating out to a Postgres instance outside Databricks, or accepting the risk until DR goes GA, and I'm curious which one teams have landed on. I'd also like to hear whether the built-in PgBouncer has held up under spikes in connections from Databricks Apps, since that's the other thing I'd want to test before putting a production workload on it. submitted by /u/InsideDebt6345 [link] [comments]

00InsideDebt63451w ago
Reddit

Laya off the benchmark: can a zero-shot decision model route real SQL traffic?

Weekend-ish experiment on Databricks. One table, TPC-H orders, 15M rows, living in two places at once: Lakebase (Databricks' managed Postgres), a continuously synced copy with a btree index on the key. Sub-100ms point lookups, useless for a GROUP BY over 15M rows. Delta behind a serverless SQL Warehouse. Great at scans and aggregations, slow at fetching one row. The synced table is the nice part: native continuous Delta to Lakebase Postgres (needs a PK + CDF), so it's one dataset under one Unity Catalog, replication handled by the platform instead of a homemade pipeline. Then I put laya (convaiinnovations/laya, a non-autoregressive zero-shot decision model, vanilla, no fine-tuning) behind a FastAPI Databricks App. It reads each query and picks the engine. MLflow traces input to decision to execution. I submitted by /u/Limp-Park7849 [link] [comments]

00Limp-Park78491w ago
Databricks CommunityLakebase Discussions

Lakebase Data API (GCP) returns jwk not found for valid service principal tokens

001w ago
Databricks CommunityGet Started Discussions

Deploying My First Databricks App with Lakebase

001w ago
HackerNews

Object Storage and WAL: Lakebase Postgres for the Agentic Era

20theanonymousone1w ago
Reddit

Lakehouse RT Questions (basics)

I see surpisingly little discussion about Lakehouse RT in this forum. Maybe databricks customers are happy using Fabric models or Duckdb as their serving layers. I'm not sure. I watched the promotional material from the 2026 summit, and learned almost nothing about the technology itself. Im assuming that it is comparable to Fabric DirectLake on OneLake, or maybe Duckdb to a lesser degree. I haven't tested yet. Pardon the basic questions... Questions - Uses SQL as query language right (no support for MDX or DAX?) -Speed is gained using tons of RAM and machine-local in-process query engines (rather than distributed MPP)? -Speed and Ram are optimized by loading from lakehouse to columnstore-in-RAM on selective columns? IE. similar to Fabric transcoding. -Query times are obviously measured from a client in the same region? RTT latency over network to an on premise user is obviously going to double the query times that matter to users. -After fresh writes to Lakehouse, the immediately following queries will be sluggish while data is reloaded to RAM? - The use of this engine with Postgres will have dramatically different performance characteristics? - Cost of using this will be higher, to reflect the large multi-core VMs needed for numerous clients, and the RAM needed to load all the related data prior to query execution. IE. this is very different than the concept of spark that runs in a distributed way on commodity hardware. - Engine is 1000% proprietary; and the dbx marketing teams wont sell this using their "open source" label like they sell everything else? - May some day have Excel plugin, like we have for Lakebase (but faster!) - May directly compete with Fabric semantic models at some point in the future Any info would be appreciated submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36972w ago
Databricks CommunityLakebase Articles

Announcement | Improving Lakebase Postgres compute cache

002w ago
Reddit

Community BrickTalk | Real-Time Data & AI: Tripwise Demo

Hey r/Databricks ! Join us for community BrickTalk on Thursday, September 24 , focusing on real-time data streaming, AI agents, and governance using Databricks. BrickTalks is a community event series where Databricks experts share real-world use cases, demos, and practical insights for building with Data and AI, giving customers a direct line to the people behind the products. In this session, we'll walk through a live demonstration of the Tripwise Demo , featuring: Sub-Second Transactions & Streaming: Device registration into Lakebase with sub-second reads/writes, plus telemetry streaming via Zerobus through a governed Medallion architecture. AI-Generated Offers & Pricing: Generating real-time agent offers using Foundation Model APIs and scoring behavioral data for usage-based renewal pricing. Natural Language Analytics: Enabling underwriters, product managers, and marketing teams to query governed insurance data in seconds using AI/BI Dashboards and Genie. Unified Governance: Managing safety, compliance, and control end-to-end with Unity Catalog. This is a great chance to see real-world architecture in action and ask questions directly to Databricks experts. When: Thursday, September 24 9:00 AM PT 12:00 PM ET 5:00 PM BST (London) 9:30 PM IST Register here and save your spot submitted by /u/Subject_Ant1789 [link] [comments]

00Subject_Ant17892w ago
Databricks CommunityLakebase Articles

Announcement | Autoscaling Lakebase Postgres

003w ago
Reddit

Unity Catalog Open Source in Name Only (UCOSINO)

Consider a callstack where something bad is happening in Spark or Unity Catalog (image above). Any software engineer will google for the message, and then for the Exception class, and then for the call frames shown on the stack (starting at the top or bottom). For any commonly encountered Exceptions from UC (something like com.databricks.sql.managedcatalog.acl.UnauthorizedAccessException), we will find dozens of results from a search engine. Others on the internet have already shared their experiences, and the search results are normally actionable. The users tell us what they had done to avoid or fix the error. But software engineers have heard for two years that "unity catalog is open source". So a software engineer will proceed to look for the source repo where they might find the full definition of "UnauthorizedAccessException", along with all the related references. No such thing exists. (Admittedly there is a public-facing github, called "unitycatalog", but it is virtually worthless and there is no overlap with the real-world UC in databricks, as we experience it.) It only takes one or two repeats of this, before a software engineer will realize that none of this stuff is actually open source. UC doesn't compare to a REAL open source software like Apach Spark. If we search for spark references in the call stack (eg. "org.apache.spark.sql.DataFrameReader"), then we are immediately taken to the source repo at github! I do give Databricks a lot of credit for open-sourcing spark. But nowadays they take too much liberty with the word "open source", to the point where it lost all of its meaning. UC is not opensource in any substantial way. Maybe there is an API spec that is open, but that is the extent of it. Another example is lakebase which the CEO claimed to be open source at the recent summit. There has never been any software as proprietary as neon/lakebase. It doesn't actually bother me if a CEO forgets how to use the term "open souce" correctly in English. What makes me more upset is when I expect to be able to use google to find the source code for "UnauthorizedAccessException", and come up with absolutely bupkis. Can anyone tell me a definition of "open source" which would potentially include either Unity Catalog or Lakebase? I'm assuming that when these words are used by the CEO, he does NOT intend to imply that the actual source is open to the public. submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36973w ago
Databricks CommunityLakebase Articles

🚀 My Lakebase App is Working!

003w ago
Reddit

I merged two databases (Postgres and Elasticsearch) into Lakebase, then threw 200 AI agents at it.

submitted by /u/Limp-Park7849 [link] [comments]

00Limp-Park78493w ago
Databricks CommunityLakebase Discussions

Lakebase Postgres branch stuck "disabled" after auto-archive → unarchive (Public Preview)

003w ago
Databricks CommunityTechnical Blog

Building Stateful Agents on Lakebase

003w ago
Reddit

What’s new in Databricks - August 2026

Databricks shipped many major Generally Available features in August 2026. Here is the breakdown of what just landed: 🚀 Unity AI Gateway Enterprise AI governance layer covering model access, Model Context Protocol (MCP) management, and cost observability. 🔒 Role-Based Access Control (RBAC) Switch to scoped, temporary role assumptions instead of dealing with permission bloat. 🔑 Secrets in Unity Catalog Unified security secrets are now governed, 3-level namespace securable objects. ⚙️ Serverless Compute Access Control Granular admin controls over who can trigger serverless workloads across your organization. ⚡ Lakebase Postgres APIs & LTAP Direct Writes Accelerated synced-table loads and improved transactional data integration. 🤖 Genie Agent Upgrades Official GA releases for both the Agent mode API and Full-page Genie Code view. submitted by /u/Youssef_Mrini [link] [comments]

00Youssef_Mrini3w ago
Reddit

Lakebase: Serverless Postgres over Open Lake Storage

Interesting paper to read about Lakebase. submitted by /u/Youssef_Mrini [link] [comments]

00Youssef_Mrini4w ago
Databricks CommunityLakebase Discussions

Lakebase Data API – First Request Times Out After 8 Seconds When Compute Is Offline

004w ago
Databricks CommunityData Engineering

Lakebase Postgre updating Delta Table.

001mo ago
HackerNews

Neon: Autoscaling Lakebase Postgres

10shenli35141mo ago
Reddit

What Is LTAP? Lakebase + Genie Explained by Databricks CTO

Want to know where data architecture is heading next? Matei Zaharia (Co-Founder & CTO of Databricks ) just broke down the future of the Lakehouse ecosystem on Here’s what he covered: 🔹 LTAP : Why real-time analytics and transaction processing are converging ? 🔹 Lakebase : The evolution of database architecture built directly on the Lakehouse 🔹 Gen ie: How AI is reshaping text-to-SQL and natural language analytics 🔹 Lakehouse RT: Unlocking ultra-low-latency real-time data streaming If you're building modern data stack architectures, this episode is a goldmine. Full video submitted by /u/Youssef_Mrini [link] [comments]

00Youssef_Mrini1mo ago
Reddit

Become Agent-Ready with Databricks Lakebase.

submitted by /u/ConstantNo2668 [link] [comments]

00ConstantNo26681mo ago
Reddit

FK Relationships and Uniqueness in UC Managed Tables

I often wish UC "managed tables" were better managed. If we wait a couple years, is it possible that relationships and unique constraints would be enforced (or perhaps validated after-the-fact)? It would be nice to have that in managed tables. It definitely seems feasible for the UC to support this in managed tables. Especially on "small" tables of a few million rows or so. Small tables would be anything that could easily be loaded into memory. If nothing else, our lakehouse-formatted storage (delta/iceberg) is good for loading a columns into memory and validating the references and/or the uniqueness of values. ... I'm guessing this is NOT going to be available in the next couple of years. That is for no other reason than we see Databricks is selling their new LTAP lakebase engine. I'm guessing that anyone who wants true referential integrity or other constraints will be redirected to this lakebase engine. Ideally there would be compelling innovations that would give reasons to stay with the "managed tables" in UC catalog. Otherwise software engineers may find reason to move to greener pastures. It doesn't make sense to validate stuff like this in custom code, when the storage engine is better positioned to do that work. submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36971mo ago
Reddit

Excel Add-In Roadmap

Has there been any public roadmap for the excel add-in? I was going to install and try it out this weekend, but I can already foresee some of its shortcomings. I'm accustomed to using Excel pivot tables for Microsoft OLAP (which are pretty hard to beat!) Questions: - Based on docs it didn't appear that this add-in would reap the benefits of lakebase (sub-ten-ms queries). Isn't that the point of using Excel, to interact with data instantly? Can we get an experience that is specifically tailored to lakebase? The CEO of databricks keeps acknowledging that "agents like fast data". But here is a newsflash; humans like fast data too! We've had fast data in Excel/SSAS pivot tables for decades. - I'm assuming this tech sends SQL queries back to the SaaS service for processing. Is that fundamentally better than the ODBC support already available to excel users? I'm guessing the catalog/usability/security is the main attraction (ie. making things "easier" and more secure). - If lakebase is as fast as the CEO claims, will it ever be possible to transpile MDX? Will those sorts of queries be able to run on lakebase? Some other open source tools do MDX, as we can see in Mondrian or Apache Kylin. These tools offer a robust, high-performance pivot table experience in Excel. Any information would be appreciated. I'm guessing it will be a very long time before Databricks wants to pursue MDX, or compete with the normal pivot tables available from Microsoft. They are more likely to follow down the current path with "metric views" for several years, rather than using pre-existing technology. From a customer perspective, I think it would be amazing if Databricks could offer an Excel experience that approaches the ones offered by Microsoft/Fabric. (One thing that is particularly compelling about the Databricks add-in is the write-back. This was something that Microsoft attempted long ago, but wasn't able to be successful with it. I'm interested to see if Databricks can do better. If nothing else, I think the culture of modern databricks user may be more receptive than the culture of the users doing write-back to OLAP cubes.) submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36971mo ago
Databricks CommunityLakebase Articles

Understanding Lakebase branching

001mo ago
Databricks CommunityLakebase Blogs

Moving Beyond Manual Changes: A Guide to Shipping Lakebase Schema with Bundles

001mo ago
Reddit

OLTP database in Databricks SaaS?

I saw the quarterly roadmap presentation. It was notable that Databricks keeps innovating with "lakebase". They simply call it their "OLTP" database offering in their SaaS. SIDE: I still feel pretty unfamiliar with this Databricks SaaS ecosystem, as compared to Fabric. Where Fabric is concerned, Microsoft has also done a similar thing. They brought their SQL Server into the boundaries of the SaaS as well, for the low-code users of that environment. In the context of Fabric, it is hard for most customers to see the point of using this "for dummies" variation of the same old OLTP database. The only scenarios for using the Fabric SQL are very contrived ... Eg . your boss makes a policy that you can use ALL the tools available in Fabric and NONE of the tools outside Fabric... even if the tools inside that SaaS are 3x more expensive than the ones outside ... and even though the ones outside the SaaS have the same 0ms network latency and the same performance. I'm still missing the vision for this lakebase OLTP offering. And it seems unusual for Databricks to start developing strategies surrounding OLTP. It seems like a very crowded space, and the only way I see Databricks being successful going down this path is if the customers are drinking one single brand of kool-aid, or else their SaaS users have some other contrived reason for not using the more affordable OLTP platforms available outside the SaaS. Can someone tell me what factors I'm missing? I admit that it is theoretically possible for lakebase to innovate and do thing that other databases CANNOT do, it seems like those innovations would only benefit 5% of customers. One example is sub-ten-ms queries out of RAM at an additional cost. Or sub-minute migration of new OLTP data to managed tables in UC catalog. If we assume that only 5% of customers might feel compelled to use this SaaS "lakebase", would that be enough adoption to allow Databricks to keep investing in this over the long term? OLTP databases have been around a LONG time, and even the smart folks at Databricks will be challenged to improve on the great and cheap options available to us! EDIT: As of a month ago, it appears that the Databricks marketing now calls it an "LTAP" database, not OLTP anymore. I'm guessing they have conceded the point that the OLTP space is crowded. I haven't yet read all the content that has been created by the Databricks marketing team; maybe that will answer all of my questions. submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36971mo ago
Databricks CommunityData Engineeringanswered

Lakebase synced table doesn’t recognize Auto CDF on a SDP materialized view

001mo ago
Databricks CommunityLakebase Articles

Learn Databricks Lakebase: managed Postgres for apps, agents & real-time data

001mo ago

Get Tuesday's version of this

Tracking Lakebase? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.

Read past issues first