Skip to content
All topics

LakeFlow

Recent items mentioning LakeFlow across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.

60 recent items4 news2 videos54 community threads

What is LakeFlow?

Lakeflow is the umbrella name for data engineering on Databricks. It bundles four pieces into one product: Lakeflow Connect for ingesting data from databases, enterprise applications, files, and message buses; Lakeflow pipelines for declarative batch and streaming transformations in SQL and Python; Lakeflow Jobs for orchestration and production monitoring; and Lakeflow Designer, a visual data preparation tool for building transformation workflows using a drag-and-drop canvas or natural language prompts.

The point is consolidation. Ingestion, pipeline logic, and scheduling used to be separate products with separate names, Delta Live Tables and Workflows among them. Lakeflow puts them behind one surface, so the connector that lands your data, the transformations that shape it, and the schedule that runs it live in one place. The pipelines layer is declarative: you state the tables you want, and the engine works out execution order and incremental processing.

Lakeflow reached general availability in June 2025. Its pipeline layer is built on Apache Spark Declarative Pipelines, a framework in the open source Spark project, and runs on the Databricks Runtime while staying interoperable with the open source version. Designer came later and entered Public Preview in April 2026.

What happened to Delta Live Tables (DLT)?

DLT was renamed and became the pipelines layer of Lakeflow. Databricks says the framework is fully backward compatible with existing DLT pipelines, so nothing needs rewriting to adopt the new capabilities.

Is Lakeflow generally available?

Yes. Databricks announced general availability on June 12, 2025, covering Connect, the pipelines layer, and Jobs. Lakeflow Designer arrived later and is in Public Preview.

Is Lakeflow Jobs the same as Databricks Workflows?

Yes, Lakeflow Jobs is the new name for Workflows, the platform's native orchestrator. At GA, Databricks said it was running over 110 million jobs per week. Databricks described the change as evolving Workflows into Lakeflow Jobs and unifying it with the rest of the data engineering stack, not as a migration.

Do Lakeflow pipelines lock me into Databricks?

The framework underneath, Apache Spark Declarative Pipelines, lives in the open source Spark project. Lakeflow pipelines are built on it and stay interoperable with it while running on the Databricks Runtime, so the Databricks-specific part is the managed runtime and platform integration rather than the pipeline model itself.

Sources: Data engineering with Databricks (Lakeflow overview), Databricks docs · Announcing the General Availability of Databricks Lakeflow, Databricks blog · Announcing the Public Preview of Lakeflow Designer, Databricks blog

What's happening in LakeFlowAI synthesis · updated 18h ago

Lakeflow Jobs recently introduced Beta support for referencing task outputs inside for_each loops 5, alongside orchestration workflows leveraging TaskValue to coordinate dynamic Python lists and arrays 12. Concurrently, practitioners are operationalizing Lakeflow Connect to handle SQL Server ingestion, schema evolution, and historical backfills 38, while using ai_query within Lakeflow jobs to query custom model serving endpoints at scale 9.

Generated daily from the 10 most recent items mentioning LakeFlow. Click any [N] to jump to the source.

Reddit

databricks TaskValue and Lakeflow

Databricks taskValues can use a Python list as input into a For each loop in Lakeflow job. You can generate the list dynamically and run the same task for every country, table, file, etc. more newshttps://databrickster.medium.com/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudekyesterday
Databricks CommunityData Engineering

Lakeflow connect SQL server ingestion can be made elastic?

002d ago
Reddit

system.query.history is now Generally Available in Databricks

Until now, answering "which query is eating our warehouse?" or "who changed that table?" meant clicking through the Query History UI one workspace at a time. Now it is a SELECT. The table logs every statement run on SQL warehouses, serverless compute, and Lakeflow pipelines, across all workspaces in the same region. Read more on LinkedIn: https://www.linkedin.com/posts/cenh_databricks-dataengineering-systemtables-ugcPost-7511094425977016320-L7uj/?utm_source=share&utm_medium=member_desktop&rcm=ACoAABmJHrsBNAC3x3H1M58JRKoHv_l4D61n0-8 submitted by /u/Lenkz [link] [comments]

00Lenkz2d ago
Reddit

Lakeflow Jobs update! Task output in foreach now available in Beta!

If you run a ForEach task in Jobs each iteration now sets a task value and a downstream task can read all of them back as on ordered array. This feature just landed in Beta - try it out! Before this, iteration outputs weren't accessible to downstream tasks. If you wanted a later task to see per-iteration results (row counts, status and output paths), you had to write each result to a table and read it back and maintain that yourself. You also lost the built-in observability Jobs gives normal task values. How it works: Inside the iteration, set a value like you always would: dbutils.jobs.taskValues.set(key="result", value=row_count) In a downstream task, read the nested task's key. You get one array across every iteration: results = dbutils.jobs.taskValues.get(taskKey="process", key="result") # results == [1, 2, 3] Or reference it as a parameter: {{tasks.process.values.result}} An iteration that never set the key keeps its slot as null: # [1, None, 3] https://i.redd.it/clkd42sm3nsh1.gif Worth knowing: Beta, needs Databricks Runtime 15.4 LTS or above Python notebooks only The assembled array caps at about 48 KB (49,344 characters) per parameter value A single parameter value can hold up to 3 aggregated references Docs: https://docs.databricks.com/aws/en/jobs/task-values#read-values-from-a-for-each-tasks-iterations And yes we are working on increasing parameter value limits! Please let us know if you have any feedback! submitted by /u/saad-the-engineer [link] [comments]

00saad-the-engineer2d ago
Reddit

Time to Swap the Cookies for Jetfuel - New Dataset & New Databricks Genie Tutorial

Most of you know samples.bakehouse . Great for a first query. Perfect for a quick demo. But after years of cookie sales, it's a little overbaked. Time to swap the cookies for jet fuel. ✈️ Together with the OpenSky Network , I brought a full day of global air traffic to Databricks Marketplace: 696 million real ADS-B position reports, messy just like real life. Myself, I used Genie for the whole journey: EDA, data exploration, a Apache Spark Declarative Pipeline, and a Lakeflow Job. Then I went a step further and read the same Marketplace data with open-source tools only, using OpenSharing and pandas. The result is this hands-on tutorial: Marketplace + Unity Catalog: get the data as a governed table Genie Agents: find anomalies in plain English Genie Agents: explore and visualize with maps and charts Genie Code: a Spark Declarative Pipeline, bronze to gold, with data quality rules Genie Code: a Lakeflow Job with schedule, retries, and email alerts Databricks Apps: your coding agent, governed by Unity Catalog OpenSharing: the open-source client in VS Code with pandas Everything runs on Databricks Free Edition (free, no credit card). 📖 Tutorial: Databricks Genie for Data Engineers and Data Scientists 💻 GitHub: databricks/tmm/DSDE-Genie-Tutorial 🛫 Dataset: OpenSky Network full-day dataset on Marketplace What's the first thing you'd query in a day of global air traffic? P.S. For the record, we still love bakehouse! 🍪❤️ [Disclaimer: I'm one of the two people who baked it.] submitted by /u/CompetitiveBet8978 [link] [comments]

00CompetitiveBet89783d ago
Databricks CommunityData Engineering

Operationalizing Lakeflow Connect:Handling Upstream Schema Evolution & Historical Backfill

001w ago
Reddit

Conditionally retrying a task on Lakeflow Jobs

I've entered a new company which uses DAB for all orchestration, and my first task has been to automate certain pipelines that fail due to data unavailability so they retry the fetch later instead of requiring manual runs. Issue is that data fetching errors are one of the many kinds that can be raised during the tasks, and it should be the only retryable one. I'm new to DAB and I've been diging into docs and forums, but I cannot find a native approach or non-patchy workaround to only retry on certain conditions (like based on exit codes). So far, I've come with these candidate solutions: - Just retrying everything (which doesn't makes sense cause if any of the other errors occur, all attempts would fail) - Failing only for the data fetch and propagating fatal errors to the following task, which would check state and fail if errors occurred - like via notebook return values (I don't like this because couples orchestration into business logic, even in a distinct layer of the one that generates the error, plus, the task that would appear as failing would be the successor to the fetching task, which would be shown as successful - not to speak about the hundreds of tasks that would need source code changes) - Failing only for the data fetch and adding an intermediate if/else task that checks if (via notebook return) the previous task propagated and error and stopping the job in such case (this is for now the most likely, as I can keep everything at orchestration level, and it is a bit more explicit even if the fetch task appears as successful, though it still feels like a workaround and tons of jobs would have to implement this new intermediate task) A couple of clarifications: - I cannot implement the retry within the task as fetches would be repeated after some time, and we don't want to have resources running and being billed just for waiting - Everything is done via Lakeflow jobs and with Databricks computer options (even if Spark is not needed) - Most of the fetching occurs by calling external FTP servers, data doesn't directly arrives to our infrastructure like to trigger the pipelines via events Thank you all Edit about explicit behavior: - On data fetch related error, retries occur - On other errors that can come from the task, the whole job fails without retrying anything - On success, next task executes - If failing due to a non retryable error, or due to maximum retry, the job fails without executing downstream tasks submitted by /u/Xandor19 [link] [comments]

00Xandor191w ago
Databricks CommunityData Engineering

How can I configure Lakeflow Connect SQL Server CDC Gateway to use a desired VM type?

001w ago
Reddit

I didn't want to choose between doing engineering and staying up to date with Databricks, so I built this. Hope it's useful to you too!

Hi everyone, I hope this post helps someone, or at least is interesting to you, because I think the problem I tried to solve is really common in our Databricks community 👀 I'm a data engineer working on Databricks (and big Databricks fan) full time. The part of the job that kept eating my evenings wasn't engineering, it was keeping up: weekly release notes, services getting renamed (a lot of you post memes and jokes about this 🥸), runtimes going out of support, and previews I'd only hear about once they'd been GA for a quarter. A lot of these "problems" turn into real problems once things are in production. Meanwhile be updated means have fast solutions before of an architectural problem will appear. And I really want to spend my time engineering, not reading updates without actually building anything. So I built a site, for everyone, that keeps all of that in one place: lakenaut.dev What's on it right now: Preview radar : features in private preview / public preview / beta, by product area. I use this to decide what not to rewrite this quarter. End of support : dates with the supported replacement next to each one. Renames : old name → new name → date. (DLT → Lakeflow Declarative Pipelines, Workflows → Lakeflow Jobs, etc.) News : what changed, who it affects, whether you need to act. A lot of the time the answer is "no", and it says so. It covers 28 areas of the platform. There's a weekly newsletter if you'd rather not check it. How it stays current: an agent reads only registered sources (release notes and docs) and proposes drafts into a review queue. I approve them. Preview stages and EOS dates get re-checked on a schedule. Not affiliated with Databricks. I just built it because I needed it, and I think it could help anyone in the same situation. Two things I'd really like to know: Which dates or previews do you track by hand today that aren't on there? Is anything wrong? There's a report link on every page, or just tell me here. I already know one big part is missing: a How-To section to help people move from an old pattern to a new one. I'm trying to figure out how to build it without losing too much of my free time (it's a personal project, and it always will be). I hope someone gets something out of it, and I'd really love to hear your opinion. Thanks in advance! submitted by /u/imthenitto [link] [comments]

00imthenitto1w ago
Databricks CommunityCommunity Articles

Real-Time SAP Accounts Receivable on Databricks Lakeflow Declarative Pipelines

001w ago
Reddit

Data Engineering Project On Real Company | Real Data | Pyspark | Databricks | Production Ready

Hello Guys, I have come across many videos on YouTube related to Data Engineering, but I found very few that actually solve a real problem. Either the problem is made up, the data is unrealistic, or the whole project is built just for the sake of showing a project. I believe that if you really want to get hands on with Data Engineering, you first need to understand the **business, the data, and the actual solution the client demands.** So, that’s what we are going to do in this series. We are going to build a complete end to end Data Engineering project, starting from understanding the business problem and going all the way to building the actual solution. We will cover things like **Lakeflow Jobs, orchestration, data pipelines, and an actual company style workflow.** The idea is simple: not just build a project, but understand **why we are building it and how it would actually work in a company.** So without any further ado, let’s begin this series. submitted by /u/Formal_Ganache3620 [link] [comments]

00Formal_Ganache36201w ago
Reddit

New features in Lakeflow Designer

New in Lakeflow Designer (screenshots below): 🚀 Genie can now create custom operators for you. For example run t-tests, create PDFs, send Slack messages etc. 🚀 Guardrails operator for Data quality checks 🚀 Multi join 🚀 Excel & CSV export controls 🚀 Output operator acts like a source for the destination it writes 🚀 Parameter suport in more operators (Combine, Limit) 🚀 Disable operators submitted by /u/zaboca_v [link] [comments]

00zaboca_v1w ago
Reddit

Anyone running actual SLOs on Lakeflow Jobs, or is it still “it failed, rerun it”?

I own the platform side for a couple of data teams. Got asked why I could produce an availability number for every service we run except the pipelines, so I went to build one off system tables. More awkward than I expected. Two things cost me an afternoon. job_run_timeline records up to an hour of runtime per row, so anything running longer is spread across several rows, and result_state / termination_code are only populated on the row that ends the run. My first success-rate query counted every intermediate row as a non-success and the numbers came out apocalyptic. Second, records land within about an hour. Fine for an error budget, useless for paging. So it ended up split: job notifications and webhooks for "wake someone up", system.lakeflow for "are we actually hitting the number". 365 days of retention is plenty for the trend. What I have now is per-job success rate and p90 duration against a target, rolled up by owner, in the same dashboards as everything else. The useful part wasn't the dashboard, it was that "this pipeline is flaky" became "this pipeline burned most of its budget in a week", which lands very differently with the people who own it. What I can't tell is whether anyone does this seriously. Most of what I read treats a job failure as a data quality thing you retry, not a reliability number someone is accountable for. So: does your org have a real target for pipeline reliability, and who owns it — platform or the data team? submitted by /u/AbilyticsEng [link] [comments]

00AbilyticsEng1w ago
Reddit

Do you still use Airflow with Databricks?

For those using Databricks in production: do you still use Airflow since Lakeflow Jobs became available? If yes, what makes you keep Airflow instead of using Lakeflow Jobs for everything? Just curious about real-world use cases. submitted by /u/almightysosa888 [link] [comments]

00almightysosa8881w ago
Reddit

Environments on classic compute

Databricks environments are my favorite way to manage external libraries. Why? Everything is in clean yml and then preinstalled on the image in the registry. Now you can use environments for classic compute as well. Just set dependency mode to environments, and then use your base environments in notebooks. more new https://databrickster.medium.com/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek1w ago
Databricks CommunityData Engineering

Predictive Optimization for Streaming tables in Lakeflow pipelines

001w ago
Reddit

Lakeflow Connect in production: What has your experience been?

I have been running Lakeflow Connect since the early gated preview (primarily testing the SQL Server connector). Overall, having ingestion natively wired into UC and smoothly moving data to silver afterwards with declarative pipelines solves a lot of headaches compared to running external tools or custom setups. That said, I have noticed some interesting trade-offs when it comes to DBU consumption on continuous syncs, edge-case schema drift, and having to resort to periodic full-refresh "fixes" when things get stuck. For those of you running it in production: How has stability and CDC replication held up for high-throughput tables? How does the total cost compare to dedicated ingestion tools like Fivetran or Qlik? Any unexpected pain points around schema evolution or operational monitoring? Has anyone had to revert to a previous ingestion method, and what did that look like? Curious to hear what is working well and where the rough edges still are. submitted by /u/OkImprovement7010 [link] [comments]

00OkImprovement70101w ago
Reddit

system tables data retention

Now you can keep data in system tables for up to 10 years, or you can set it to clean up faster. more news https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek2w ago
Databricks CommunityBrickTalks TV

BrickTalk Recording | One Platform, Any Source: Unifying Enterprise Data with Lakeflow Connect

002w ago
Databricks CommunityData Engineering

Lakeflow connect CT pipeline keeps running multiple connections in source

002w ago
Reddit

[Discussion] Lakeflow Jobs: If your data catalog could trigger a job on any change what would it be?

In addition to triggering on time (cron) intervals you can trigger on data landing: files arriving on a share, table gets a new commit and (soon) jobs completing upstream. While we were building Model update triggers (which trigger when UC registered models change), we were discussing the possibility of triggering on any entity in UC’s metadata changing. Examples here could be: Someone tags a column PII and you want to trigger a scan or retention job Someone changes a grant and an access review job gets triggered Someone is given SELECT permission on a table More generally it could be: Schema change on an upstream table (column added, dropped, renamed, type changed): run compatibility tests, rebuild the downstream model, ping the owner before it breaks. PII or sensitivity tag added: kick off a scan, apply retention, open an access review. Grant or ownership change: access recertification, sync to the entitlement system, log for audit. Table created, renamed or dropped in a schema: auto-register or clean up downstream assets. Questions: Which of these would you use and what are just noise? What kind of change would you really really want that you currently do by hand or after the fact? Anyone doing this off audit logs, system tables or just polling running jobs? submitted by /u/saad-the-engineer [link] [comments]

00saad-the-engineer2w ago
Reddit

Private Network Gateway

Private Network Gateway is one of the year's biggest network innovations. Serverless can now be part of your VNET! more news https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek2w ago
Databricks CommunityData Engineering

Best practices for SLA monitoring and automated retries across hundreds of Lakeflow Jobs

002w ago
Databricks CommunityData Engineering

igrating from Cron-based Airflow to Lakeflow's Data-Aware Triggers — Real-world experiences?

002w ago
Reddit

Lakeflow Connect SQL Server Connector

I recently enabled Lakeflow Connect (lfc) on the source database - the issue is, some of the tables in the source database (managed by another team) does NOT have a primary key (which means that in lfc, a __databricks_id is used to identify a unique record). Thus, the DBAs enabled CDC on the source database. However, when I ingested the data into DBX using the Lakeflow Connect Managed SQL Server Connector, one of the tables in the source database had duplicate records (two or more records with the same value across all columns). This caused my Lakeflow Connect pipeline to break. Any ideas on how to fix this? (Other than dropping duplicate records in the source DB and implementing a unique constraint on the source DB)? I was wondering if there is a specific setting in Lakeflow Connect that I can toggle that I'm missing. submitted by /u/RazzmatazzLiving1323 [link] [comments]

00RazzmatazzLiving13232w ago
Databricks CommunityTechnical Blog

Tutorial: Transform your Lakeflow Connect ad data into visual and conversational analytics

003w ago
Reddit

Read this if you use Streaming Tables in Lakeflow Spark Declarative Pipelines

🚀 We’re excited to announce that Lakeflow Spark Declarative Pipelines (SDP) now supports creating “vanilla” (i.e., non STREAMING) MANAGED TABLES and writing to them via one or more append flows , using the new CREATE TABLE ... FLOW ( SQL ) and create_table() (Python) APIs . What is this Beta? This Beta allows creating a managed table that is populated by append flows: CREATE TABLE ... FLOW (SQL) / create_table() + @append_flow (Python) create a managed table written by one or more flows. Fan multiple sources into one table — declare several flows targeting the same managed table. Full table surface works: partitioning, liquid clustering, expectations, row filters, table properties, and private (pipeline-local) tables. import_checkpoint on append_flow , which migrates an existing Structured Streaming workload into a pipeline without reprocessing the source — the flow imports the query's existing checkpoint and resumes from the last committed offset with state intact. Example (Python): from pyspark import pipelines as dp dp.create_table("combined") dp.append_flow(target="combined") def from_a(): return spark.readStream.table("source_a") u/dp.append_flow(target="combined") def from_b(): return spark.readStream.table("source_b") Example (SQL): CREATE TABLE events PARTITIONED BY (bucket) FLOW INSERT BY NAME SELECT id, bucket FROM STREAM read_files('abfss://my_path', format => 'json'); Where do we need help? We are in Beta, so there might be some rough edges. Please take this for a spin and share your feedback here . What’s next? Managed Tables support for other flow types (AutoCDC, Replace Using, and Replace Where) is coming soon! Learn more CREATE TABLE ... FLOW (SQL reference) — https://docs.databricks.com/aws/en/ldp/developer/ldp-sql-ref-create-table-flow create_table (Python reference) — https://docs.databricks.com/aws/en/ldp/developer/ldp-python-ref-create-table import_checkpoint on append_flow — https://docs.databricks.com/aws/en/ldp/developer/ldp-python-ref-append-flow Questions, feedback, or help: comment below or share feedback in the form: https://forms.gle/7bGP5FYN7P1Z4WP27 submitted by /u/SlightImagination250 [link] [comments]

00SlightImagination2503w ago
Reddit

UC secrets in Key Vault

Secrets in Unity Catalog is a great feature introduced a few weeks ago, but since then, everyone has been asking to use Azure Key Vault as a secrets backend. Thanks to rapid development, we can now link our schema to Azure Key Vault; UC will read secrets as UC secrets, and permission management will be through Unity Catalog. In that scenario, you insert/update secrets in Azure Key Vault, but read/reference and grants can go through UC. more news https://databrickster.medium.com/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek3w ago
Reddit

Google Drive connector in Lakeflow Connect is now generally available (GA)

The Lakeflow Connect connector for Google Drive is now generally available ! It’s now easier than ever to ingest structured and unstructured files from Google Drive into Delta tables for analytics and AI workloads. You can configure a managed ingestion pipeline through the UI or managed API. Managed pipelines automatically handle incremental processing, automatic retries with exponential backoff for source API rate limits, failure recovery, and provide rich Google Drive metadata. For direct control over ingestion logic, you can also just use the Spark + SQL APIs directly: spark.read , Auto Loader, read_files, or COPY INTO pointed at Google Drive URLs. https://preview.redd.it/j65yaaft3uoh1.png?width=2048&format=png&auto=webp&s=0f9f13cc63572b991232a8fa617aa1fe6697369c Link to public docs + references: Google Drive managed connector documentation Spark + SQL APIs and examples Community blog and video tutorial: From PDF to insights Data + AI Summit session: Intelligent Document Processing with Lakeflow Common workloads include: Loading Google Sheets, Excels, CSV, JSON, and other structured files into Delta tables. Ingesting PDFs, Google Docs, Google Slides, and images. Parsing documents with ai_parse_document to prepare content for extraction, search, and agents. Examples of using the Spark + SQL APIs: Read an Excel sheet from Google Drive with spark.read : ​ df = (spark.read .format("excel") .option("databricks.connection", "my_gdrive_conn") .load("https://docs.google.com/spreadsheets/d/9k8j7i6f...")) Ingest unstructured documents + PDFs from a Google Drive URL with read_files , then easily parse them using ai_parse_document : ​ CREATE OR REFRESH STREAMING TABLE gdrive_documents_table AS SELECT *, "_metadata" FROM STREAM read_files( "https://drive.google.com/drive/folders/1a2b3c4d...", format => "binaryFile", `databricks.connection` => "my_gdrive_conn", pathGlobFilter => "*.{pdf,docx}"); CREATE OR REFRESH STREAMING TABLE documents_parsed AS SELECT *, ai_parse_document(content, map('version', '2.0')) AS parsed_content FROM STREAM gdrive_documents_table; Coming soon: Ingest Google Drive’s per-file permissions and ACL metadata to power permission-aware AI agents, enterprise search, and more. If you try it, share what you are building and let us know if you hit any friction! submitted by /u/BricksterJ [link] [comments]

00BricksterJ3w ago
Reddit

SharePoint connector in Lakeflow Connect is now generally available (GA)

The Lakeflow Connect connector for Microsoft SharePoint is now generally available! It’s now easier than ever to ingest structured and unstructured files from SharePoint into Delta tables for analytics and AI workloads. You can configure a managed ingestion pipeline through the UI or managed API. Managed pipelines automatically handle incremental processing, automatic retries with exponential backoff for source API rate limits, failure recovery, and provide rich SharePoint metadata. Soon, our managed connectors will also support ingesting SharePoint Lists and per-file permissions metadata. For direct control over ingestion logic, you can also just use the Spark + SQL APIs directly: spark.read , Auto Loader, read_files, or COPY INTO pointed at SharePoint URLs. Common workloads include: Loading Excel, CSV, JSON, and other structured files into Delta tables. Ingesting PDFs, Word documents, PowerPoint files, and images. Parsing documents with ai_parse_document to prepare content for extraction, search, and agents. https://preview.redd.it/89i379aattoh1.png?width=2180&format=png&auto=webp&s=292350dfddfc9bc1a4daf3fc4447394821206054 Link to public docs + references: SharePoint managed connector documentation Spark + SQL APIs and examples Community blog and video tutorial: From PDF to insights Data + AI Summit session: Intelligent Document Processing with Lakeflow Examples of using the Spark + SQL APIs (after first creating a UC connection ) : Read an Excel sheet from SharePoint with spark.read : excel_df = (spark.read .format("excel") .option("databricks.connection", "my_sharepoint_conn") .option("headerRows", 1) .option("dataAddress", "Sheet1!A1:M20") .load(" https://mytenant.sharepoint.com/sites/Finance/Shared%20Documents/Monthly/Report-Oct.xlsx") ) Ingest unstructured documents + PDFs from a SharePoint URL with read_files , then easily parse them using ai_parse_document CREATE OR REFRESH STREAMING TABLE sharepoint_documents_table AS SELECT , "_metadata" FROM STREAM read_files( " https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents ", format => "binaryFile", databricks.connection => "my_sharepoint_conn", pathGlobFilter => " .{pdf,docx}"); CREATE OR REFRESH STREAMING TABLE documents_parsed AS SELECT *, ai_parse_document(content, map('version', '2.0')) AS parsed_content FROM STREAM sharepoint_documents_table; Coming soon: Ingest SharePoint Lists into Delta tables (coming super super soon) Ingest SharePoint’s per-file permissions and ACL metadata to power permission-aware AI agents, enterprise search, and more. If you try it, share what you are ingesting and where you hit friction! Don't hesitate to ask questions! submitted by /u/BricksterJ [link] [comments]

00BricksterJ3w ago
Reddit

[Discussion] Lakeflow Jobs: How do you use table update triggers?

(databricks product manager here) Curious how people are using table update triggers in production. https://docs.databricks.com/aws/en/jobs/trigger-table-update Do you rely on the available debouncing capabilities (protect against over and under triggering) or do the options feel confusing enough that you mostly work around them? When a trigger needs to represent more than “run when this table changes,” how do you express the business logic? For example: Do you use control or checkpoint tables to signal that an upstream workflow has finished? Do you wait for a specific status, batch ID, watermark or set of tables before starting downstream work? Do you put that logic in the trigger itself, or in a separate workflow/job? What has worked well and what has been difficult to reason about or debug? Any other suggestions or feature requests relating to Table Update Triggers? I’m especially interested in real-world patterns and whether the current debouncing behavior is intuitive enough for you, or whether a control-table pattern ends up being the clearer approach. Edit: how many folks still use control tables instead of data tables with these triggers? Thank you 🙏 https://preview.redd.it/munx0r2u4ooh1.png?width=660&format=png&auto=webp&s=80e6b80a62cf77870bf444e7634d1ee7412feee3 submitted by /u/saad-the-engineer [link] [comments]

00saad-the-engineer3w ago
Reddit

No more UNION ALL-ing all of your SDP pipeline event log tables for monitoring

https://preview.redd.it/3w0j4ss3inoh1.jpg?width=2048&format=pjpg&auto=webp&s=50bc79280a546544e8f795947e5af6b211e9596a Hi folks, Databricks PM here - I wanted to share an exciting update that you no longer have to manually publish and combine your pipeline event logs for monitoring across pipelines and workspaces. We just launched the beta for the pipeline events system table (system.lakeflow_pipeline_events_preview.pipeline_events). Key features: All pipeline events (regardless of cluster start) are automatically captured without any manual enablement or maintenance. Events are aggregated in a central system table without requiring any custom aggregation logic. This data is available at close to real time latency (based on our internal testing we achieve a P99 latency of less than 1 minute). An admin can grant a single user access and they can query events for every pipeline in the table. Fine-grained access controls support which scopes the visibility to only the pipelines the user has access to is coming soon. The data remains in the system table even after pipeline deletion and is retained for 13 months. If you want longer retention this is also possible with the configurable retention feature for System Tables. Query ergonomics are better now with the use of VARIANT. Here are some sample queries in case you want to try them out: -- The latest error for each pipeline that has errored in the last 7 days, with the outermost exception. -- The exception chain is ordered with the root cause last, so read element -1 for the root cause. -- On many errors only the first element carries error_class and sql_state. SELECT workspace_id, pipeline_id, event_time, event_type, message, error.exceptions[0].error_class AS exception_error_class, error.exceptions[0].sql_state AS exception_sql_state FROM system.lakeflow_pipeline_events_preview.pipeline_events WHERE level = 'ERROR' AND event_time >= current_timestamp() - INTERVAL 7 DAYS QUALIFY ROW_NUMBER() OVER (PARTITION BY workspace_id, pipeline_id ORDER BY event_time DESC) = 1 ORDER BY event_time DESC -- Flow throughput for a specific pipeline SELECT origin.flow_name, date_trunc('HOUR', event_time) AS hour, SUM(variant_get(details, '$.flow_progress.metrics.num_output_rows', 'BIGINT')) AS rows_written FROM system.lakeflow_pipeline_events_preview.pipeline_events WHERE pipeline_id = ' ' AND event_type = 'flow_progress' AND event_time >= current_timestamp() - INTERVAL 7 DAYS GROUP BY origin.flow_name, date_trunc('HOUR', event_time) ORDER BY hour DESC, rows_written DESC -- Data quality: failed expectations by dataset, per update, in the last 1 day SELECT pipeline_id, update_id, origin.dataset_name, expectation.name AS expectation_name, SUM(expectation.failed_records) AS failed_records FROM system.lakeflow_pipeline_events_preview.pipeline_events LATERAL VIEW explode(variant_get(details, '$.flow_progress.data_quality.expectations', 'ARRAY >')) AS expectation WHERE event_type = 'flow_progress' AND event_time >= current_timestamp() - INTERVAL 1 DAY GROUP BY pipeline_id, update_id, origin.dataset_name, expectation.name HAVING SUM(expectation.failed_records) > 0 ORDER BY failed_records DESC; Beyond single queries you can build alerting (using Databricks SQL alerts) and dashboards. We will share a new dashboard template soon - I will update this post once its available. Call outs: This is in beta right now, if you are not opted in we will not capture your event log data. Enablement: Toggle on the “ Lakeflow Pipeline Events System Table ” from the account level preview. Docs are linked here , would love to hear your thoughts on how you will use it or what else you want to see to improve observability! submitted by /u/brickster_123 [link] [comments]

00brickster_1233w ago
Reddit

How to organize your notebook tabs?

It was a real pain, but now, with a few tricks, you can manage them better. First, in Workspace files, next to the DABs folder or git repo, there is a small shortcut to show only tabs from that DABs folder or git repo. Alternatively, you can also use the switcher in Home next to Notebook. If you need to organize your tabs differently, there is new functionality: spaces, which let you group them however you like. more news https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek3w ago
Reddit

Cost-optimized way to reflect source DB changes in Silver in <1 minute?

Due to new business requirements, we need to reflect the state of a few source DB tables (5 to 40 million rows each) in the Databricks Silver layer in less than 1 minute. Currently, the flow looks like this: Source DB → AWS DMS in CDC mode (ingests new data every 30 seconds to S3) → S3 landing bucket → DLT pipeline running on serverless compute in continuous mode. The DLT pipeline ingests the append-only data into the Bronze layer using file notification mode and updates the Silver layer using an Auto CDC flow. This works great, and we achieved what we wanted with relatively low effort because we already had DMS in place. We just added an extra replication task to ingest data more frequently for the tables we need. However, in this setup, the DLT pipeline costs are quite high. Ingesting just 6 Bronze tables and 6 Silver (Auto CDC) tables costs around $50 per day, which is about $1,500 per month. For comparison, DMS, which replicates more than 800 tables to S3, costs us less than half of that. My question is : is there any other more cost-optimized option we could consider to achieve less than 1 minute latency when reflecting the source DB state in the Silver layer? Maybe Lakeflow Connect or some custom process? Extra notes: - I know that adding more tables to the DLT pipeline makes the cost per table lower because Databricks can optimize the clusters more efficiently. - I know that using a cron schedule could reduce costs, but for these particular tables, we can’t use a schedule like every 10 minutes or similar because we need the data to be updated in less than 1 minute. - I know that for the relatively small tables currently in scope, we could eliminate the Auto CDC flow and create a normal view on top of the Bronze table, with deduplication and deletion logic. This would slightly sacrifice query performance, but we expect more similar use cases in the future, so I’m looking for a solution that can scale. submitted by /u/CyberEnzo [link] [comments]

00CyberEnzo3w ago
Databricks CommunityCommunity Articles

Learn Databricks Lakeflow | Ingest, Orchestrate, and Build pipelines on one platform.

003w ago
Databricks CommunityData Engineering

Using Databricks Asset Bundles and Lakeflow Jobs in a Real Project

003w ago
Reddit

Community BrickTalk | One Platform, Any Source: Unifying Enterprise Data with Lakeflow Connect

Hey r/Databricks ! We’re hosting a free, community-sponsored BrickTalk on Thursday, September 17, 2026, focusing on how to simplify and scale data ingestion using Lakeflow Connect! BrickTalks is a community event series where Databricks experts share real-world use cases, live demos, and practical insights, giving you a direct line to the people building the products. Stop struggling with fragmented data across disparate sources. In this session, we'll demonstrate how Lakeflow Connect enables seamless data ingestion from SaaS apps, databases, and cloud storage directly into the Databricks Platform with zero infrastructure management. 🛠️ What We’ll Cover Native Data Ingestion: Learn how Lakeflow Connect provides fully managed ingestion directly into Unity Catalog as governed Delta tables. Simple Integration: See how to easily connect data sources using a simple UI or API. Accelerated AI & Analytics: Discover how unifying your data powers Customer 360, Operations, and downstream AI agent workloads. ⏱️ Global Times PT: 9:00 AM ET: 12:00 PM BST (London): 5:00 PM IST: 9:30 PM 👉 Register here to save your spot! submitted by /u/Subject_Ant1789 [link] [comments]

00Subject_Ant17893w ago
Databricks CommunityAnnouncements

Community BrickTalk | One Platform, Any Source: Unifying Enterprise Data with Lakeflow Connect

003w ago
Databricks CommunityGet Started Discussions

Tech Companies & Lakeflow Connect

003w ago
Databricks CommunityData Engineering

Lakeflow SDP Append Flow

003w ago
Reddit

Serverless Env v6

Version 6 of the serverless environment is available, which corresponds to runtime 19. more news https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek3w ago
Databricks CommunityData Engineering

Lakeflow connect Ingestion pipeline notification for gateway pipeline

004w ago
Databricks CommunityTechnical Blog

Lakeflow Connect: Message Bus Ingestion - Now shipping your logs directly! (Beta)

004w ago
Reddit

Lakeflow Genie Code Task

We now have Genie Code Task in Lakeflow jobs. Can not yet send output to if/else, but more options for orchestration are planned. more news https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek4w ago
Reddit

Dynamic Select is now in Lakeflow Designer

You can now dynamically select columns in Lakeflow Designer. This makes it easy to bulk-keep, or bulk-drop columns from a very wide table. And your data prep will keep working as your underlying schema evolves. submitted by /u/zaboca_v [link] [comments]

00zaboca_v4w ago
Reddit

External price control

In the Unity AI gateway, it is also possible to register an external model for which we pay the provider directly (OpenAI, Anthropic, etc.). In that case, Databricks now knows the prices for those models and can calculate, monitor usage, and alert or block based on budgets. more news https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek1mo ago
Reddit

Lakeflow Designer - Source and Output operators now support parameters

You can now use parameters in Source and Output operators for Lakeflow Designer. This is pretty useful for iterating on your Data Prep in dev/test before running it in prod. submitted by /u/zaboca_v [link] [comments]

00zaboca_v1mo ago
Databricks CommunityCommunity Articles

All 18 Lakeflow AUTO CDC configurations went green. Five failed my ship check

001mo ago
Reddit

SDP-Meta Deep-Dive Demo: Building Data Pipelines at Scale on Databricks (w/ Databricks Sr. Staff FDE)

This s a helpful resource for those building pipelines at scale on Databricks. From the docs: SDP-META is a metadata-driven framework for Lakeflow Spark Declarative Pipelines . Define your Bronze and Silver pipelines in a JSON or YAML onboarding file — a single generic Declarative Pipeline reads the resulting DataflowSpec at runtime and builds the full processing graph automatically. No pipeline code to write. Who it's for: platform and data engineering teams standardizing repeatable Bronze/Silver pipelines across many datasets — onboarding new feeds through metadata instead of new pipeline code, with consistent data quality, quarantine, CDC, clustering, and sink patterns available through Bundles, CLI, UI, MCP, and agent workflows. When it's not the best fit: one or two simple pipelines, Gold-layer business modeling, tables that each need unique application logic, a managed connector and downstream logic that already satisfy the complete Bronze/Silver requirement, or a need for a formal support SLA (SDP-META is a Databricks Labs project). See the Introduction for the full positioning. You can find the project at https://github.com/databrickslabs/sdp-meta submitted by /u/JosueBogran [link] [comments]

00JosueBogran1mo ago
Reddit

Skills in Unity Catalog

Skills are available in Unity Catalog. They use a similar concept to volumes and are integrated with the AI gateway. New REST endpoints for skills are coming, and a new tool to manage them, ucode, is already available. more news: https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek1mo ago
Databricks CommunityData Engineering

Best practices for data quality in lakeflow

001mo ago
Databricks CommunityData Engineering

Lakeflow connect

001mo ago
Reddit

Bye Bye Fivetran

I went to the Ingestion section and saw that Fivetran is no longer there. It was always there for many years. Also, at the same time, a few new Lakeflow connectors were added. submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek1mo ago
Reddit

Build up a data history in Databricks based on Azure SQL data

I have the following need: there's an OLTP Azure SQL DB, which holds transactional data for a period of roughly 30 days only. Now that data should be replicated to Databricks delta tables with a maximum delay of 15 minutes, not as a 1:1 copy, but instead growing over time. Ideally the timeframe covered on Databricks side should be several years. From what I've read so far, either CDC or CT with Lakeflow should be the way to go. The only thing I'm worrying about are breaking schema changes: as this is an OLTP DB managed by a different team, we have no chance to prevent such as incompatible column type changes (e.g. from a string to a date type), column renames or even column drops. I thought about using Views managed by the other team instead, as some kind of an abstraction contract, but neither CDC nor CT are applicable on Views. How did you guys solve such a requirement? Would also appreciate to hear some best practices of Databricks consultants based on real customer solutions. submitted by /u/Sea_Basil_6501 [link] [comments]

00Sea_Basil_65011mo ago

Get Tuesday's version of this

Tracking LakeFlow? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.

Read past issues first