Ce que la communauté demande.
Discussions récentes de r/databricks et du tag [databricks] de Stack Overflow : problèmes concrets, questions d'intégration et cas particuliers à connaître.
This week
10 questionsProblems with jobs / GIT repo
Virtual Event | The Emerging Blueprint for Agentic Apps
dhqodh odw qwd
djqwgiuld q qoiwdh iqw hdoiqw
Build Your First Databricks App with Codex: A Practical Setup Guide
How to choose the right AI agent path on Databricks
Data Lakehouse with Agentic AIs: A Guide
submitted by /u/codingdecently [link] [comments]
Data Lakehouse with Agentic AIs: A Guide
submitted by /u/codingdecently [link] [comments]
Genie is so Dumb and I am tired of Pretending Otherwise
Hey Guys, Data Engineer with 3.2 YOE here, I have been working on and off on databricks for multiple projects across domains like Retail, logistics and Now Pharma. In my current role, I focus on creating end to end pipeline with different data sources which often get refreshed on a bi-weekly basis for pharma related data The client is extremely demanding and has no technical expertise to gauge how much time it really takes to build and maintain something so complex and therefore expect the team to use Claude teams license and also Genie The problem with Genie is that, I can not trust it, It will not listen to my instructions even after having a separate instructions.md which I update on a daily basis, in fact I update this after each session is complete, and no, I do not use AI to update this, every change request that the client and the client's team has goes through my own words of instruction updates, I have an entire markdown file which has contexts for multiple islands of projects, workspace folders and notebooks that I have to maintain and transition to devops team. I also have created 4 different skill family for Genie, in relation to migration, devops, maintainence and review. And it is so disgustingly bad at handling all this context, that it fails me across all 4 areas. I have seen it lie to my face multiple times, Casting columns as null, altering data types without asking consent, never following my work stream and process in a sequential format Having set the wrong notebook paths in a task even though i explicitly tell it exactly what to do. It often confuses streams of works and ends up mixing so many things that remove this knot and fuck up itself costs me so much time!!! It is very poor in adhering to the domain specific instructions that we provide More than 3 requests in a chat, and it is practically useless And branching of chats is the most useless feature that they have introduced. I have created multiple diagnosis queries to check the work of this agent , and it will straight up lie to you, so much, so convincingly, that you know it knows all the biases you have and it even goes a step ahead and just narrates a story that you can provide to your team in the stand-up. My only concern is, after all this bullshit, my team, of nearly 6 (some of them have never worked in databricks or delta table environments like this) was charged $990 dollars in the month of august , for this shit output??? Have some shame Databricks , fix your worthless product or remove that feature or stop charging so much if you are beta testing in actual prod. Today, I am writing this post because of something that I caught live, that pushed me to the edge. Something that is critical, production level issue, which If I was not paying attention while merging could have been a huge disaster, and mind you the pipelines I make are client facing, what is even more distrubing is the fact that it will just make so many unnecessary changes to a simple query or a pyspark function just enough so that the tests are passed. (So basically it does not want to get caught, and makes a mistake so that we can prompt it again to fix this mistake and then Databricks can charge us more, this looks like a dark pattern to me) In your preview (prima-facia), everything is good, the tests are passing, the job is running, but holy!!!, it was casting 3 columns which are essential for downstream processes as null. It boldly suggests that we remove those 3 columns, and or cast them as null. How is this a solution Databricks? The AI is supposed to have more context than me because it is a agent made specifically for Databricks Environment Correct?? That is how it is sold?? And upon spending just 2 mins, I realized that the fix that it is suggesting will blow up critical information that the client team should see in the app because it is not even there in the first place, and all of this is because it can not read the correct notebooks , even if you tag it. It had se […truncated]
Is Databricks Classic Compute getting too heavy for small workloads?
I’ve been using Databricks Classic Compute for a while, and recently I’ve started wondering whether cluster cold starts are getting noticeably heavier with newer DBR versions. For example, with DBR 18 LTS, I tried a small 2-vCPU VM for a single-node job cluster and hit DriverStartupTimeout after 300 seconds. Databricks even suggests that this commonly happens on instances with fewer than 4 CPU cores. That feels a bit surprising for workloads that are not actually Spark-heavy — e.g. running Python, dbt-core, API calls, or using Databricks mainly as a job runner inside a VNet. I like Classic Compute because of the flexibility and straightforward VNet/private networking. Serverless is attractive for startup time, but in our environment it would mean quite a bit more networking setup. So I’m curious: Have you noticed Classic Compute cold starts getting slower or more resource-hungry across newer DBR versions? Do you now consider 4 vCPUs the practical minimum for a reliable driver? Has anyone benchmarked the same VM size across DBR 14/15/16/17/18? What are you doing for lightweight non-Spark workloads where you still want Classic Compute? Update: Single node DBR 18 LTS + Standard_D4pls_v6 cold start spent almost 11 minutes 'Waiting for resources' ( Azure Japan East ) https://preview.redd.it/ta90ftommmth1.png?width=442&format=png&auto=webp&s=9815bf14c15362cfe526e96e7904557da11bb839 submitted by /u/bobjia-in-tokyo [link] [comments]
Last week
116 questionsI built an open-source check that stops AI agents from inventing column names in SQL.
Databricks ai_decide explained in 5 minutes!
submitted by /u/ConstantNo2668 [link] [comments]
App Spaces: governance for apps
Workspace admins, thanks to App Spaces, can define who can create and use apps in a given space and also set policies for apps, for example, which API scopes can be used. more news https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]
Exams with additional fees are included?
Where do i start if i want to learn databricks concepts. I want to get start using my learnings to apply at work.
submitted by /u/Time_Meringue_1300 [link] [comments]
Ingest data from Microsoft Outlook
Databricks ai_decide: The AI Judge That Never Writes a Word
Use Databricks ai_decide() to grade every ai_query output in SQL and decide what gets sent, without a second LLM submitted by /u/Lenkz [link] [comments]
My data saved in the account lost
Databricks SDP (Spark declarative Pipelins) Overwrite table
Millions in Vacuumable files, but only 10-20 GB recovery size
Hello Folks. Wanted your opinion on something - I am seeing numerous cases of our tables having millions of redundant vacuumable files, but total recoverabke capacity is only 10-20 GB . Is it worth Vacuuming them? My sense is that the DBU and ADLS listing charges will be more than cost of that storage. Is there any instance where Vacuuming will pay off. Also, if deciding to Vacuum - what is better, Job Compute or Serverless SQL Small Size. Thanks. submitted by /u/Lumiaman88 [link] [comments]
Built a live Helsinki tram tracker on Databricks with Lakeflow, Lakebase, Genie and Apps
Certificate Not Received After Passing Databricks Certification Exam
Cross-Catalog Sync: Iceberg on Polaris, Glue, and Unity
submitted by /u/codingdecently [link] [comments]
Building A Data App with dashboard and agentic capabilities in dbx
Hi all, I’m currently working on a POC and im given the flexibility by my boss to explore using any tool (im considering mainly databricks or power apps now) to build out a “dynamic” data app to replace powerbi . I am fully aware that dbx ai/bi dash does not has nearly 100% capabilities of that the power bi dash but I’m just trying to showcase a different experience of reporting that leverages AI and not POWERBI (boring! to leaderships). For more context, im working with SAM/CMDB EoSL data that is all sitting in dbx unity catalog already. My expectations is something simple that looks like a dashboard but i can utilize genie to communicate pr generate any sql findings within the dashboard context. Something easy to initiate and maintain in the future. Additionally, it would also be best that the views created that caters to this app can be reused as native queries to powerbi in case the need in the future that id still have to migrate it there. End product expectations: The end product is expected to be easily accessible without any additional access required for any user that wants to access the app or genie. Think of leaderships and executives using the dash mainly. It should be something that is easy to access and managed within a workspace like what powerbi serves currently which is what all the stakeholders are currently used to already. I am no expert in Databricks but have been an active user for about a year now. All the best to everyone, cant wait to hear what you guys have in mind in the comments! (this is my first reddit post btw) submitted by /u/Junior-Confusion-691 [link] [comments]
How are you attributing agent cost per successful task, not just per token?
[[Guía@Delta-México]] PAso a PAso ¿Cómo hablar con Delta en México?
Let AI Decide: SQL Decision Making with Databricks ai_decide()
Databricks just introduced ai_decide(), a new AI function now available in Beta. Give it text and your questions, and it returns a category, a probability, or a score directly in SQL. Use those results to drive decisions in your workflows: route requests, prioritize work, or determine when human review is needed. submitted by /u/hubert-dudek [link] [comments]
PowerBI refresh disabled
Exams with additional fees are included?
Omnigent 0.16.0 + Hermes on macOS: "hermes did not accept the message" and .env/auth.json copied into every session. Anyone fixed this?
I'm new to this, so apologies if I'm missing something obvious. Setup: macOS (Apple Silicon), Omnigent 0.16.0 installed with uv tool install omnigent (Python 3.12), Node 22, tmux 3.7, Hermes Agent installed via git (up to date). Running fully local on 127.0.0.1:6767. Problem 1: first message gets dropped. When I start a new Hermes session from the Omnigent desktop app, it fails with: inner executor error: hermes did not accept the message (the TUI may still be initializing); no new transcript row appeared after two delivery attempts then Required terminal exited unexpectedly; the session runtime is no longer available. The logs show Hermes is still starting up (installing dependencies, browser tool checks) when Omnigent pastes the message, and then Hermes exits with a KeyboardInterrupt. Running omnigent hermes in one-shot mode from the terminal works fine. Problem 2: credentials copied per session. Every Hermes session creates a new folder under the macOS temp dir ( .../T/omnigent-501/hermes-native/ /hermes_home/ ) with a copy of my Hermes .env and auth.json , plus about 1–2 GB of runtime. After one afternoon I had 19 folders (22 GB) and about 30 copies of my keys sitting in temp. Questions: Has anyone gotten the Omnigent desktop app to reliably start Hermes sessions? Is there a setting to give Hermes more startup time? Is there a way to make Omnigent use an existing Hermes profile without copying its .env / auth.json each session (for example, a shared HERMES_HOME or a secrets manager)? Is either issue fixed in a newer version, or is there a GitHub issue I should follow? I'd like to eventually run a dedicated Hermes profile (one with sensitive API keys) through Omnigent, so the credential copying is the big blocker for me. Thanks! submitted by /u/Kwontum7 [link] [comments]
Rescheduled my exam of My Data Brick ML Engineer Associate Exam
Premium PAYG account stuck at TRIAL_VERIFIED (rate limit of 0 on resold models) — how do I get promo
How do you guys handle schema changes in Databricks when the source keeps changing?
Like the source suddenly adds a new column, changes a column type, or removes something. Do you make the pipeline handle these changes automatically, or do you just let it fail and fix it? Iam.. curious what approach actually works better in real projects, especially when the source keeps changing. submitted by /u/Bhanuprakash_1947 [link] [comments]
Analyze
Lakeflow connect
Noxivam – En enkel kapselrutin för vuxnas dagliga välbefinnande
Serverless compute is not available in my account/workspaces
Postgres is now the most efficient search database at scale
How to fit ~1GB+ embedding model into a 2Gi Kubernetes pod? Getting OOMKilled
Hi, Deploying a FastAPI to Kubernetes that uses a multilingual sentence-transformers embedding model (ONNX backend, CPU only). My source data lives in a Delta table in Databricks, and the app reads from it and generates embeddings. The app takes user input (text) at embeds it, and compares it against stored embeddings for semantic similarity. So at least the query embedding has to happen live. Pod resources: - CPU: 1 request / 2 limit - Memory: 2Gi request = 2Gi limit (platform policy requires memory request and limit to be 1:1) Current docker setup: - Multi-stage Docker build (python:3.11-slim) - CPU-only PyTorch - The model is downloaded at build time, so it's baked into the image Image breakdown: - HF model cache: ~1.1 GB - torch: ~650 MB - pyarrow, scipy, transformers, pandas: ~100–150 MB each - Plus sklearn, onnxruntime, mlflow, and others The pod gets OOMKilled at startup or shortly after. With a ~1GB model, torch, pandas/pyarrow, and the ONNX runtime session all in one process, I think I'm just over 2Gi. What I'm trying to figure out on where do you store large models in production? Would love to hear what setups have worked for you. Thanks! submitted by /u/runningnozone [link] [comments]
Feature Request: pipeline_task should have "no wait" option
Sometimes we want to trigger a pipeline (such as database sync) without having to wait for its completion. We can still subscribe to notifications on the pipeline itself to react to an unexpected error, so the pattern should be fine. submitted by /u/CarelessApplication2 [link] [comments]
The trendtrack promo code is BESTCOUPON20
Resource limits
I am trying to build an interactive dashboard on the underlying Databricks. Which of these are the best?
I’m thinking of 4 options here 1. Build an MCP (a custom MCP) that can access custom tools on Databricks and interface it on Claude.ai or Claude desktop Advantage- Claude is very good at inferencing, multi-turn conversation and multi-step processing Disadvantage - custom MCP and tools needs to be built accurately and validated. It should have full context of schema and unity catalog Build a semi-custom MCP - this will use “askGenie “ as one of its tools with additional custom tools Advantage- complexity decreases as we leverage genie space Disadvantage- double inference by genie and Claude Use custom Databricks connector in Claude. Not sure if this uses genie and therefore double inference but it’s more reliable than custom build because this is a native offering by vendor Use only genie and build custom dashboard without needing Claude interface What’s the thought on this? submitted by /u/Bala_Devaraj [link] [comments]
databricks TaskValue and Lakeflow
Databricks taskValues can use a Python list as input into a For each loop in Lakeflow job. You can generate the list dynamically and run the same task for every country, table, file, etc. more newshttps://databrickster.medium.com/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]
PARTITIONED BY, Liquid Clustering, and OPTIMIZE: when should you use each?
Join the Live AI/BI + Genie Launchpad (Oct 6–8)
CUSTOMER STORY | BookMyShow scales self-service analytics with Genie
The Breakdown: Databricks
submitted by /u/Roadtochessmaster [link] [comments]
'DATABRICKS DEN' - NEW COMMUNITY
Over the past few weeks, I've been building something special. Every day I speak with Databricks experts across Europe. After so many of those conversations, the obvious question was: why not build a community around it? So today I'm launching 'Databricks Den'. Whether you're looking for your next contract or just want to stay close to what's happening in the Databricks market, the Den is for you. Join now to catch the first round of insights: 🔗 Link to the Den Looking forward to seeing you there 👊 submitted by /u/Reuben_UMATR [link] [comments]
Gemini Foundation Model endpoint unavailable in 14-day commercial trial
Routing between GLM 5.3 Flash and GLM 5.3 with ai_decide: a cheap router in front of ai_query
Document AI is a pipeline, not a model
Setting my Expectations for Microsoft CSS support in Azure Databricks (2026)
I recently moved back to using Azure Databricks after a multi-year hiatus. And I opened my first support ticket this week. My prior support experiences were long ago (in 2021). I don't remember these support experiences being particularly terrible. But nowadays there are some Microsoft platforms where their support work is subcontracted to remote organizations like Mindtree-India. These organizations (one or more of them) may provide the first, and second tiers of support before any FTE from Microsoft takes the reins. Is that how things will work in the context of Azure Databricks? If/when the FTE at Microsoft is unable to help with a ticket, how long will it take for the support to be transferred from Azure Databricks to Databricks? I have a sev B ticket on "standard" support at the moment and I thought a human would engage within a few hours but that isn't happening. I will probably just give up for today. Can anyone offer tips about how to navigate support tickets in Azure? Any advice would be greatly appreciated, and I think it could make the difference between a ticket that lasts two days or two weeks. Thanks. submitted by /u/SmallAd3697 [link] [comments]
Budget override option not available
The Four Architectures That Make AI Work
submitted by /u/Berserk_l_ [link] [comments]
Row-level security for a RAG agent: what Unity Catalog enforces, and what you have to build
Lakeflow connect SQL server ingestion can be made elastic?
MLflow traces accepted by StartTraceV3 (200 OK) but never stored: GetTrace returns NOT_FOUND
system.query.history is now Generally Available in Databricks
Until now, answering "which query is eating our warehouse?" or "who changed that table?" meant clicking through the Query History UI one workspace at a time. Now it is a SELECT. The table logs every statement run on SQL warehouses, serverless compute, and Lakeflow pipelines, across all workspaces in the same region. Read more on LinkedIn: https://www.linkedin.com/posts/cenh_databricks-dataengineering-systemtables-ugcPost-7511094425977016320-L7uj/?utm_source=share&utm_medium=member_desktop&rcm=ACoAABmJHrsBNAC3x3H1M58JRKoHv_l4D61n0-8 submitted by /u/Lenkz [link] [comments]
How to build FinOps chargeback by cost centre — allocating the whole cloud bill, not just your DBUs
The 32nd Column: Why Your Z-Order or Clustering Key May Be Doing Nothing
Databricks Lakebase Search - GA Announcements
submitted by /u/Wise_Ear_4064 [link] [comments]
Announcement | Five ways marketers can use Genie One
At what data size do you stop reaching for Spark?
Photon enabled but a large share of the plan is falling back, cost up and runtime flat
Streaming read fails with "Detected a data update" after a restatement job touches the source
Unable to connect to Tableau Cloud
Serverless - how is this OK for a "fully managed" product?
Notebook that's run fine for months suddenly started failing with 'No Module Found 'chardet' Nothing changed on our end. Seems that serverless environment v6 dropped chardet from the base requirements, and since the notebook wasn't pinned, the "default" just rolled forward underneath it. No warning, and the release notes don't even mention the removal. I get the argument of pin your environment, declare your dependencies. But serverless is sold as the "no config, we manage it for you" option. Silently removing packages from the default environment mid-week, with zero notice, feels like a breaking change shipped as a patch? submitted by /u/OneSeaworthiness8294 [link] [comments]
Lakeflow Jobs update! Task output in foreach now available in Beta!
If you run a ForEach task in Jobs each iteration now sets a task value and a downstream task can read all of them back as on ordered array. This feature just landed in Beta - try it out! Before this, iteration outputs weren't accessible to downstream tasks. If you wanted a later task to see per-iteration results (row counts, status and output paths), you had to write each result to a table and read it back and maintain that yourself. You also lost the built-in observability Jobs gives normal task values. How it works: Inside the iteration, set a value like you always would: dbutils.jobs.taskValues.set(key="result", value=row_count) In a downstream task, read the nested task's key. You get one array across every iteration: results = dbutils.jobs.taskValues.get(taskKey="process", key="result") # results == [1, 2, 3] Or reference it as a parameter: {{tasks.process.values.result}} An iteration that never set the key keeps its slot as null: # [1, None, 3] https://i.redd.it/clkd42sm3nsh1.gif Worth knowing: Beta, needs Databricks Runtime 15.4 LTS or above Python notebooks only The assembled array caps at about 48 KB (49,344 characters) per parameter value A single parameter value can hold up to 3 aggregated references Docs: https://docs.databricks.com/aws/en/jobs/task-values#read-values-from-a-for-each-tasks-iterations And yes we are working on increasing parameter value limits! Please let us know if you have any feedback! submitted by /u/saad-the-engineer [link] [comments]
CUSTOMER STORY | ANA Seguros turns insurance data into faster decisions with AI/BI and Genie
Quick clarification about GENAI or Context engineering certification
Step-by-Step: Building a Vacation Rental Operations App with AppKit
Azure Databricks, two-week-old account: partner pay-per-token endpoints (Claude, Grok) have never once succeeded — "Databricks-set rate limit of 0" — and the Assistant has "no daily token allowance". Open models work. What gates this?
Azure Databricks, Premium, one account with three workspaces, Unity Catalog, Azure Marketplace billing (no card, no trial credits left). The account was bootstrapped on a trial workspace about two weeks ago and moved to Premium six days later. I've spent two days on this and want a sanity check from anyone who has seen it. The facts, from system.serving.endpoint_usage: • Every open-weight pay-per-token endpoint has worked since the trial: gpt-oss-120b has ~50 successful calls going back to the trial period, Llama 3.3 70B likewise, zero refusals. • No partner pay-per-token endpoint has ever returned a success on this account. The very first call anyone made to databricks-claude-sonnet-5 was refused, and every call since, on every Claude endpoint and on Grok, in all three workspaces, for users and service principals alike: 403 PERMISSION_DENIED: The endpoint is temporarily disabled due to a Databricks-set rate limit of 0. • It is not our AI Gateway config: I removed every rate limit from one Claude endpoint and invoked it — same 403. The endpoints' config shows nothing else. • The notebook Assistant worked during the trial. On the paid account it refuses everyone, admins included: *"This workspace has no daily token allowance for the assistant."* Underneath, /ajax-api/2.0/conversation/llmproxy/ returns 429 {"type":"daily_token_limit_reached","daily_limit_tokens":0,"message":"…Daily token limit of 0 reached. Resets at midnight UTC. (DTB)"} and it recomputes to 0 every midnight. • Once, the Databricks Apps in the workspace were stopped by the platform with *"App compute was stopped due to workspace or account status"* while every workspace showed RUNNING. apps start brought them back. Red herring, for completeness: we also created a Unity AI Gateway per-user budget whose first version had a $0 threshold with BLOCK_USAGE. That blocked Genie for a day, exactly as documented, and was fixed. It was created a day *after* the first Claude refusal, so it isn't the cause of any of the above. What I've found: two Community threads describe "Databricks-set rate limit of 0" as a workspace trust-tier gate — trial-born accounts sit in TRIAL_VERIFIED, pay-per-token partner models are gated to PAYABLE_VERIFIED, a card alone doesn't flip it, only Databricks Sales/Support moves the tier. Both were AWS/personal accounts, neither resolved on-page. The account console's own API reports our account's feature_tier as STANDARD_W_SEC_TIER although the workspaces are Premium; I can't find what that field means. Questions • Does Azure Databricks with Marketplace billing go through the same PAYABLE_VERIFIED gate? Does it clear on its own after the first settled invoice, or does someone have to move it? • If you were moved: who did you contact (account team, or "Contact us" in the console) and how long did it take? • Is the Assistant's "daily token allowance" the same gate, or a trial allowance that simply goes to 0 when the trial ends on an unverified account? • Does feature_tier: STANDARD_W_SEC_TIER on the account object mean anything to anyone? submitted by /u/sumit671 [link] [comments]
Built my first Databricks App with Lakebase
Dashboard DAB Deployment: default_catalog and default_schema do not work for metric views anymore
Let's talk about AI rate limits
Databricks Unity Gateway is great and getting even better very quickly. It now has nearly everything we could hope for. New frontier models are usually available within 24 hours of release. But...the models aren't really available . They all come with extremely low rate limits, way too low to support a team of engineers in any sizable enterprise using the models. If you have a sizable agents.md even 1 person hits the limits. Why is it like this? Sure we can raise a ticket to request higher limits but that takes weeks and the highest we've gotten is 1M tok/min. We can (and do) go to Bedrock and get 5M+ tok/min by default, no special request needed. Is the expectation that we hook up other providers? Are the key frontier labs (OpenAI, Anthropic, Google) limiting Databricks as a whole? For such a great product this seems like a glaring limitation that's holding us back. Make it make sense. submitted by /u/degenbets [link] [comments]
DB as a SIEM
My organization is considering replacing our current SIEM with DB. Wondering of anyone has done it and looking for feedback. Specifically looking for any gotchas, architecture recommendations. Cheers! submitted by /u/Environmental_Arm370 [link] [comments]
Spark and observability
submitted by /u/ZenithR9 [link] [comments]
Writing to externally managed Iceberg tables from Databricks: am I missing something?
I recently joined a new company as a data architect and I'm trying to understand how things work here today (new to Databricks, main experience with RedShift and BigQuery), and I'd like corrections from people who use Databricks daily. I've been testing how Databricks handles Iceberg tables that are managed by Glue. What I'm seeing: Can read tables as Databricks' federation lets me mount the external catalog and query the tables as foreign tables. Can't write (is this accurate?) INSERT, MERGE, and UPDATE operations against those foreign Iceberg tables fail, and the Databricks docs seem to confirm this. Question: Has anyone gotten writes to non-UC Iceberg tables working from Databricks? That could be through a config, a preview feature, or an open-source Spark Iceberg runtime with a REST catalog on a cluster. If I've got any of this wrong, please correct me as I am still getting up to speed with the platform. submitted by /u/DonutSignificant4174 [link] [comments]
DevOps Fundamentals: Lecture - DevOps Fundamentals: Repeated word
1.Not able to access interent (gcp databricks) 2.both ingress and egress access given
Tip: Why your Delta MERGE fails with "multiple source rows matched" (and the simple fix)
Azure to AWS
The lakehouse serving fight is on
Serverless Interactive and Realtime Inferencing costs scaling rapidly
Hello folks, I have recently started a new role as Data Platforms leader, and while going through our monthly Databricks billings - two things stood out, both mentioned in Title, costs for them scaling rapidly month on month. Serverless Interactive compute, by default has access for Admins inherited, and people now dont want to go back to the slow time taking All Purpose Compute starts. Serverless Realtime Inferencing has me confused, as I cannot see any such compute - is that cost for using Genie? Is yes, can Genie access be turned off for everyone in the Workspace? Just these two items reduced will save 20-25% costs for us. Thanks in advance for any pointers. submitted by /u/Lumiaman88 [link] [comments]
Can a Delta table be “correct” even when the source and target row counts match?
Suppose the source and target both contain 100M rows. The row counts match, but the target could still have: Missing records Duplicate records Incorrectly transformed values Changed or truncated values For a large Delta table, what checks would you use beyond row counts to prove that the source and target data actually match? Would you use hashes, key-level reconciliation, aggregates, sampling, or something else? submitted by /u/No_Ambition8323 [link] [comments]
Implementing a "Zero-Bus" Architecture with Unity Catalog
Implementing a "Zero-Bus" Architecture in a Lakehouse: Best Practices for Shared Dimensions?
Announcement | The Genie One MCP is now Generally Available
How to make each person see only their specific data (i.e. their own rows in Genie) with RBAC & ABAC
Tag automation with databricks
No table description? Not queried for 90 days? See how to automatically tag the table and notify its owner in my new article. https://medium.com/databrickscommunity/clean-up-the-mess-with-tag-automation-72ae119b6f35 https://www.sunnydata.ai/blog/databricks-unity-catalog-automated-tagging submitted by /u/hubert-dudek [link] [comments]
Solution Accelerator Series | Incident Investigation Using Graphistry
Lessons learned from configuring Genie Agents
Certification Support Request #01015865 – No Response Since September 14
Data Quality in Microsoft Fabric – Native features vs. external tools (like Databricks Expectations)?
Hey everyone! I’ve been working a lot with data quality frameworks in Databricks (leveraging things like Delta Live Tables expectations or custom validation notebooks), and I'm curious about how the community is handling Data Quality in Microsoft Fabric. Does Fabric have a native, built-in feature for data quality/expectations (similar to what Databricks offers), or are most of you relying on external tools, libraries (like Great Expectations/Soda), or custom PySpark/SQL validation notebooks inside Data Engineering pipelines? How are you currently enforcing quality checks before data hits your Gold/Semantic layers in Fabric? Would love to hear your approaches and best practices! I read this documentation here: https://learn.microsoft.com/en-us/purview/unified-catalog-data-quality-fabric-lakehouse?WT.mc_id=510336 , but it doesn't say much. submitted by /u/SuperbNews2050 [link] [comments]
WORKSPACE HAS NO DAILY TOKEN ALLOWANCE
Deployment Jobs structured yaml.
Time to Swap the Cookies for Jetfuel - New Dataset & New Databricks Genie Tutorial
Most of you know samples.bakehouse . Great for a first query. Perfect for a quick demo. But after years of cookie sales, it's a little overbaked. Time to swap the cookies for jet fuel. ✈️ Together with the OpenSky Network , I brought a full day of global air traffic to Databricks Marketplace: 696 million real ADS-B position reports, messy just like real life. Myself, I used Genie for the whole journey: EDA, data exploration, a Apache Spark Declarative Pipeline, and a Lakeflow Job. Then I went a step further and read the same Marketplace data with open-source tools only, using OpenSharing and pandas. The result is this hands-on tutorial: Marketplace + Unity Catalog: get the data as a governed table Genie Agents: find anomalies in plain English Genie Agents: explore and visualize with maps and charts Genie Code: a Spark Declarative Pipeline, bronze to gold, with data quality rules Genie Code: a Lakeflow Job with schedule, retries, and email alerts Databricks Apps: your coding agent, governed by Unity Catalog OpenSharing: the open-source client in VS Code with pandas Everything runs on Databricks Free Edition (free, no credit card). 📖 Tutorial: Databricks Genie for Data Engineers and Data Scientists 💻 GitHub: databricks/tmm/DSDE-Genie-Tutorial 🛫 Dataset: OpenSky Network full-day dataset on Marketplace What's the first thing you'd query in a day of global air traffic? P.S. For the record, we still love bakehouse! 🍪❤️ [Disclaimer: I'm one of the two people who baked it.] submitted by /u/CompetitiveBet8978 [link] [comments]
How to reduce query latency?
How to validate delta tables in Databricks?
im working with delta tables in databricks and wanted to know how ppl usually validate the data after an etl job runs. apart from checking if the job completed, do we have to check things like record counts, duplicate, nulls etc Like if the source has around 100k records, how do you make sure the expected records reached the Delta table? for incremental loads, do you also check that only the expected records were added or updated? And how do you handle schema changes, like a new column being added or a data type changing? Do you build these checks directly in databricks itself? Would be interested to know what checks ppl actually use in production. submitted by /u/Klutzy_Solid5200 [link] [comments]
How to bring the project id info into the databricks billing usage table on GCP
Avoid Unity Catalog external location permission bugs by using managed storage
Hey, I work as a Databricks Engineer at Abilytics, so here's my take. If you are migrating existing tables to Unity Catalog, defining external locations incorrectly will break your grants. When you register an external location via CREATE EXTERNAL LOCATION, Unity Catalog validates your storage credential against the underlying cloud storage container. However, users often forget that read and write access on the external location is distinct from the IAM role or service principal trust policy. If a user has SELECT on a table, but lacks BROWSE on the external location, queries against parquet files directly via path-based access will fail with permission denied, even if the table-level grant succeeded. Always verify your storage credentials have explicit container-level policies attached before mapping external volumes. Furthermore, remember that path-based access like spark.read.load('s3://my-bucket/path') requires explicit external location grants, whereas managed tables abstract this entirely. Prefer managed tables default storage roots to bypass manual external location ACL synchronization issues entirely. Happy to go further — more on Databricks | Platform Engineering | AI at Abilytics , AI Systems, Data Engineering & Platform Engineering Services. submitted by /u/AbilyticsEng [link] [comments]
Stop using standard worker nodes for intermittent batch jobs
Hello, I'm part of the Databricks Engineer team. If you are running short-lived batch pipelines on standard clusters, you are wasting money on idle instance termination times. Switch your job clusters to use spot instances by configuring the instance pool or setting the spot bid policy directly within the Databricks Jobs API or UI cluster settings. When configuring a job cluster, set the spot instance policy to use spot with fallback to on-demand to prevent job failures when spot capacity is scarce. For intermittent batch jobs and maintenance tasks like OPTIMIZE and VACUUM, using spot instances significantly reduces compute costs compared to persistent multi-node on-demand clusters. Ensure your retry logic handles node preemptions gracefully, and leverage job clusters rather than all-purpose clusters so the cluster terminates immediately after the batch execution completes, eliminating idle billing. If this is useful, there's more on Databricks | Platform Engineering | AI at Abilytics — AI Systems, Data Engineering & Platform Engineering Services. submitted by /u/AbilyticsEng [link] [comments]
Courses still showing as in progress
Data fixes on a databricks schema. Thoughts?
Hi all, I have a databricks schema of a few hundred delta tables that need some data fixes for specific records in each of those tables . This schema itself is raw data and gets ingested into some downstream data tables and the fixes have been requested by business. Now, importantly, this schema is the most upstream source for this data - the SQL db it was ingested from no longer exists. I had thought about doing the fixes via a transformation layer in the pipelines that ingest this raw data but given how many tables there are with records that need updating I don't really want to create hundreds of new 'data fixed' tables. I also read that it's best to fix something as upstream as possible. Given these are delta tables, rolling back should theoretically be possible if something goes wrong. Anyway, that's my rationale for making changes to the source tables. My question is what your preferred method is to make data fixes? I obviously need something where its easy to rollback if needed. I can obviously just achieve this with python migration scripts and use delta timetravel in case something goes wrong, but wonder if there are recommended libraries or tools for the job that have what I need out of the box? Should I also keep an unchanged copy of this schema or is that too redundant? submitted by /u/Spooked_DE [link] [comments]
Databricks Micro Apps, App Spaces and Genie App Generator
Databricks launched App Spaces and Serverless Micro Apps at the end of last week (in Beta). Over the weekend, I tested migrating two of my production Databricks Apps to App Spaces. App Migration 1: Blocked by Zero Egress My Data Portfolio Project Creator required internet access to pull data stacks from live job postings. Then, the LLM needs internet access to research for open data sources to use. In standard Databricks Apps, this runs cleanly. In App Spaces, there is zero external internet egress including for LLMs. The app can reach internal workspace resources, but it cannot touch the outside web. If your app relies on third-party APIs, external databases, or web scraping, App Spaces is a non-starter until Databricks opens network egress. Workload 2: Success on Internal FinOps My client-facing DBU cost observability app reads workspace usage data and writes directly to Lakebase. Because it requires zero external network calls, the migration worked. Both the micro app compute and Lakebase scale to zero when idle. Cold starts take roughly 30 seconds (though in beta, you occasionally need a quick browser refresh once it spins up). For internal, low-frequency administrative tools, this turns a continuous monthly compute bill into pennies. (+ App Spaces, Genie App Generator, and Micro Apps are free while in beta) The next part is less about App Spaces and more just general best practice for Databricks Apps that I see people miss. Stop Using Delta Lake as an OLTP Database Databricks Apps are software applications, not batch analytics notebooks. If your app writes application state, session data, or row-level CRUD directly into analytical Delta tables, you need to rethink that design. For app transactions, use Lakebase (serverless Postgres which also scales to 0). Both are governed under UC, but Lakebase gives your app the low-latency transactional engine that application engineering actually requires. My Verdict on the Beta App Spaces solves the idle compute problem that has plagued Databricks Apps since launch. But until Databricks allows us to deploy directly from existing Git repos and opens external network egress, it remains limited to internal-only use cases. Curious on other peoples experience... How has it been for others? submitted by /u/OkImprovement7010 [link] [comments]
Business data context is the key to better answers.
submitted by /u/ConstantNo2668 [link] [comments]
Declarative Pipeline Expectations - Quarantining Bad Records
No compute fleets in Azure Databricks?
I see that AWS has supported spot-priced compute fleets for a few years. But Azure still doesn't use fleets to select spot-priced VM's (yet). Customers are forced to set up our own spot-priced pools and manage a fleet of them independently. Two years ago someone named Walter (an FTE at databricks?) said "currently on development but no ETA " https://community.databricks.com/t5/administration-architecture/compute-fleets-on-azure-databricks/td-p/104754 So what is the reason for so much delay in Azure? Is there some sort of conspiracy going on where Microsoft and Databricks are teaming up to give us only expensive compute (like serverless) and restricting our ability to access other alternatives? This is getting a bit frustrating. submitted by /u/SmallAd3697 [link] [comments]
Salesforce Hosted MCP + UC HTTP connection — stops working after ~1 hour?
Workshop on bringing systematic evaluation and MLflow tracking to LLM apps, Oct 3
If you're already using MLflow for experiment tracking on the ML side, there's a decent chance your LLM work hasn't caught up to the same standard yet. Most teams are still hand-tuning prompts and eyeballing outputs while everything else in the stack gets versioned and tracked properly. Serj Smorodinsky and Brett Kennedy, co-authors of a book on LLM applications, are running a live 3-hour workshop on Oct 3 that applies that same rigor to LLM development. What's covered: Programming LLM behavior with DSPy signatures and modules instead of hand-written prompt strings Building a baseline classifier live, from a real task Constructing an evaluation dataset with task-specific metrics so "better" is measurable, not a feeling Reading failure patterns directly out of the eval results Few-shot and instruction-level optimization applied systematically on top of the baseline Experiment tracking and LLM trace management through MLflow, so every run is reproducible and comparable Saving and reusing optimized DSPy programs across projects Communicating LLM reliability to stakeholders, which usually gets skipped entirely It's basically the MLflow discipline this community already applies to models, extended to prompts and LLM pipelines. Aimed at people already past "which model should I use" and into "how do I make this reliable." Full details here. submitted by /u/camerongreen95 [link] [comments]
How to get people to actually choose Genie: train the mindset, not just the buttons
DataSmart
Claude Academy's Using Databricks for Data Analysis - This official documentation from Anthropic on Claude Academy for how to work with Databricks is really underwhelming, so I wanted to share a guide that properly explains everything.
Hey Databricks community! I have been working with Databricks as an admin for over 4 years and I was attempting to follow this official tutorial from Anthropic for how to connect Claude with Databricks https://academy.claude.com/tutorials/using-databricks-for-data-analysis And I encountered multiple significant frustrations with doing it, as the tutorial is very underexplained and outdated for how to get Claude working with Genie One and other functions within Databricks. Therefore, I thought this video on Youtube would be valuable content to share to help solve this problem, for other people that are encountering similar difficulties when getting these two tools connected and building out systems that work for Claude to properly interact with Databricks. Let me know if this video is helpful! submitted by /u/k_kool_ruler [link] [comments]
Connecting to Snowflake via Workload Identity Federation
Announcement | EBOOK: Operationalizing Agent Fleets at Scale
Your exam is being suspended for non-compliance with requirement set by test sponsor
SDP pipeline parameters
Recently attempted to use pipeline parameters with my SDPs but it seems they are only supported in SQL at the moment. This is a PITA for me using python as I am using the pipeline config as an alternative but this doesn't support updates at runtime. Could you confirm if pipeline parameters are coming to python SDPs and if so, when? submitted by /u/dvartanian [link] [comments]
Change data feed from a materialized view
Your Databricks query is slow — what would you fix first?
Imagine you have a Delta table with several billion rows . A query that used to take 30 seconds now takes 8–10 minutes. The query itself looks reasonable, but the table has: A large number of small files Data that is not well organized for the most common filters Several joins Frequent incremental writes You could approach the problem in several ways: OPTIMIZE the table Change the partitioning strategy Use liquid clustering Improve the query itself Change the join strategy Reduce the amount of data being scanned Reconsider how the data is written What would you investigate first, and why? Would you start with the Spark UI/query plan, the table's physical layout, or the SQL itself? I'm especially interested in different approaches rather than one specific solution. submitted by /u/No_Ambition8323 [link] [comments]
Knowledge Assistant fails to load after indexing; generated AI Search index cannot be reused
Community BrickTalk | Databricks Certification: Did You Know?, What's New, and Where We're Going
What Are the Key Considerations for Building AI-Powered Applications?
Liquid clustering vs. partitioning for a 5 TB Silver table with frequent MERGEs
What information about external agents can be retrieved from Unity Catalog?
Unlock the Power: Databricks Genie One vs. Code and Agent
Week of Sep 21
74 questionsGenie App Builder
What apps are you guys building? submitted by /u/Sea-Glass7015 [link] [comments]
trigger on_bundle_deploy
In a job bundle, we can specify the trigger on_bundle_deploy in job_runs. Once the bundle is deployed, the job will run automatically. This is really useful for post-deployment tasks like DDLs or one-time table population, such as a calendar. More news and the best articles https://medium.com/databrickscommunity submitted by /u/hubert-dudek [link] [comments]
Preparing for Databricks Champion — looking for guidance
I’m planning to prepare for the Databricks Champion program and wanted to connect with people who have already gone through the process. submitted by /u/Ok-Golf2549 [link] [comments]
Architecting a Medallion Lakehouse
Balancing Model Agility and Centralized Governance with Mosaic AI Gateway
UC Berkeley renames football stadium to Databricks Field in sponsorship deal
FTA: Cal Athletics has announced a multiyear partnership with Databricks, a data and AI platform for enterprises, that includes renaming Cal football’s home field to “Databricks Field at California Memorial Stadium.” The San Francisco-based company was founded at UC Berkeley and now will have its name attached to a wide variety of aspects of Cal football, most notably the field at the stadium and the helmets of the players. Sounds like Databricks didn't have to pay any cash upfront for the rights, but is instead giving UC Berkeley stock in exchange. While Cal Athletics and Learfield did not disclose the financial terms or length of the partnership — besides stating that the multiyear partnership will be funded through UC Berkeley becoming a shareholder in Databricks — the deal is reportedly worth $22.8 million for up to 10 years, according to Ben Portnoy of FOX Sports. submitted by /u/gamescan [link] [comments]
Jobs suddenly stopped running with no errors
Data Lakehouse Architecture Guide
submitted by /u/codingdecently [link] [comments]
Use Genie to crate a dashboard to show cost per catalog or workflow
Hi there ! Is is feasible to ask Genie and guide me to create a way of showing databricks costs (we used to do it in azure by adding labels). We want to have it directly in databricks. We use catalogs for each client. Each client has a workflow. Please let me know your thoughts or if you have better suggestions. submitted by /u/francebased [link] [comments]
Zempira: dagelijkse capsules voor darmbalans en gewichtsbeheer
AUTO CDC SCD Type 2 with late-arriving deletes: edge cases the docs don't cover
Stop Moving Data Into Spreadsheets. Databricks Is Moving the Spreadsheet to the Data.
submitted by /u/TowardsDataExp [link] [comments]
Accidentally removed the only workspace admin – Free Edition
Databricks now deletes unused service principal secrets after 90 days
(removed)
Accidentally removed the only workspace admin – Free Edition
Customerlake - first impressions, pointers, highlights?
I'm about to dive in, and wanting to set my expectations. Advice from real people appreciated. submitted by /u/KeyMammoth1348 [link] [comments]
Agents in Databricks
Hello! For some time I have been thinking in AI Engineering and I was wondering what kind of agents I could deploy in Databricks with the entire AI module. Have you implemented agents in PRD? What tasks or value do these agents add to their processes? A while ago I had the idea of creating a PII classification agent, until Data Classification has been released. I would like to hear your experiences and get inspired. Thanks! submitted by /u/Weekly_Marionberry_3 [link] [comments]
questions on sdp-meta
Planning on adopting sdp-meta , i have few questions. i may test drive it in couple of days but want to check with users who may have adopted it. If we are consuming 100's of tables daily dump via AUTOLOAD , Do i infer schema or should i derive schema for individual table Can i apply any custom logic transformations from bronze to silver other than simple select expr and filters like group by and joins. submitted by /u/Competitive-Fee-4006 [link] [comments]
Open-source app on Databricks system tables - looking for people to build it with me
Declaring a Primary Key RELY turned my dbt unique Test Green on Duplicate Data
Operational Responsibilities in DBX
Has anyone come across job responsibilities that aren't software developers (engineers, analysts, or scientists)? In other words, I'm wondering if there is such a thing that corresponds to a classic DBA position, but for this Databricks SaaS? Now that UC is here with managed catalogs, the technology is looking more like a conventional database platform. In fact the new lakebase is actually a conventional Oltp based on postgres (or LTAP if you prefer the latest marketing). It seems like the ecosystem is complex enough now to support some operational job positions, especially if the data estate is large. Anyone have a position that is more than 50 pct operational?) What about 100 pct operational, like a devops or DBA that focuses on databricks submitted by /u/SmallAd3697 [link] [comments]
CIF architecture diagram inconsistent with the description
Is Databricks Replacing Power BI and Tableau in the Enterprise?
Why did Databricks get into the BI tooling space & what does it mean for Databricks customers using Power BI, Tableau, and other established BI tools? In this 3-min video, David Meyer (SVP @ Databicks) shares how Databricks can help both provide advance decision making capabilities and lower costs, as well as where some of the older tools fit in your stack. This was recorded in front of a live audience at the Databricks Data + AI Summit 2026. Hope you enjoy it, and would personally love to hear your stories, both where AI/BI Dashboards from Databricks has been able to take on more of your BI workloads, as well where you still find gaps that need to be addressed! submitted by /u/JosueBogran [link] [comments]
Learn Databricks AI Agents | Build, Deploy & Evaluate on Databricks Platform
Best Architecture Approach for Databricks Multi-Cloud and Multi-Region Deployment
Configuring Genie: practical steps for more reliable answers
Databricks Workflow Task Failure - Custom Error Messages
Tech summit FY 2027
Databricks Usage in real world?
Hi all, I want to understand how AZure Databricks is actually being used in real world data engineering teams. For those of you working with Databricks, how do you use it in your current job? Likee, do you write most of your code in notebooks and publish those notebooks through Git? Or are you using Asset Bundles for ci/cd and deploying jobs and workflows through different environments? I’m trying to get a clear picture of how ADB is used in actual companies. Help this buddy out. TIA!!! submitted by /u/Wise_Beat_7035 [link] [comments]
DeepSeek V4.1 Flash returns workspace ITPM rate-limit error on fresh/idle workspace
Genie One Foundations: What Data Teams Should Get Right Before Rolling It Out
Bricktalk Recording | Real-Time Data & AI: Tripwise Demo
Azure devops yml deployment of asset bundle
I need help on deploying asset bundles from Azure Devops. Where does the azure-pipelines.yml live, what parameters are required in the azure-pipelines.yml, does anyone have a basic example of it working as I'm completely lost in the complexity of it. Thanks. submitted by /u/GardenShedster [link] [comments]
org.apache.iceberg.connect.IcebergSinkConnector to sink data from kafka to databricks - large volume
Demo Material: Building and Orchestrating an Agentic App - AI Agent Fundamentals Plan
Where to find Notebooks for course 'Advanced Techniques with Apache Spark Declarative Pipelines'
Anyone migrated from legacy CDF to Auto CDF yet?
Lakebase handles failover within a region, but cross-region disaster recovery is still in Private Preview. How are people covering that?
I was reading through Databricks' breakdown of what Lakebase takes over on the ops side. Patching, scaling, in-region failover, and point-in-time restore with 2 to 30 days of history all run automatically, while cross-region disaster recovery is Private Preview on AWS only, with manual failover and recovery procedures the customer owns. For anyone running Lakebase behind an app or for agent state, how are you handling a region outage today? I can see people copying scheduled snapshots elsewhere, replicating out to a Postgres instance outside Databricks, or accepting the risk until DR goes GA, and I'm curious which one teams have landed on. I'd also like to hear whether the built-in PgBouncer has held up under spikes in connections from Databricks Apps, since that's the other thing I'd want to test before putting a production workload on it. submitted by /u/InsideDebt6345 [link] [comments]
Managed Disaster Recovery
Do you actually partition your Delta tables, or has everyone moved away from it?
keep seeing conflicting advice around partitioning Delta tables, and I'm curious what people are actually doing in production. I understand the general idea of partitioning, but in practice I'm finding it harder to decide when it's genuinely useful versus when I'm just creating more small files / directories for the sake of it. For example, imagine a table with hundreds of millions of records that's mostly queried by date, but also has filters on things like customer, region, status, etc. Would you partition by date? Leave it unpartitioned and rely on Delta's optimizations? Something else? And for people who have actually dealt with this in production — what made you change your original approach? I'm particularly interested in “we did X, it caused Y problem, so we changed it to Z” stories. Those are usually more useful than the textbook explanation. submitted by /u/Delulu62134 [link] [comments]
USE CONNECTION bypasses your MCP Service's Tool Selection and Policies
UC ingestion from a SQL server
I'm creating a UC catalog. Lets say I want to ingest a SQL database that contains 50 tables, parented by five meaningful SQL Server schemas. (Fact.WhateverThing, Dim.WhateverThing, Logging.Whatever, CONFIG.Whatever, and VIEWS.Whatever). I plan to name the catalog in a conventional way ("sales_dev",) but I'm stumped for naming UC schemas and UC tables. Since the Dim and Fact tables need to be joined frequently, I'm assuming all these tables should be in a SINGLE unity catalog schema, yes? Should I simply move all five of these SQL schemas into a single UC schema (maybe into a single UC schema named "bronze_from_sql")? At that point how would I name the tables? I'm assuming I would be forced to do something like so: "fact_whatever_thing", "dim_whatever_thing", "logging_whatever", "config_whatever", "views_whatever"? It strikes me that UC doesn't have the same concept of schema that we have in SQL server. A schema in UC is actually referring to a whole database. Whereas in SQL it can be used to organize tables within a database. Edit: decided to publish sql server tables to volumes (in a bronze layer). Will just use delta format, and free-form folder-and-table names. Is very liberating not to conform to the constraints of the structured table names in UC submitted by /u/SmallAd3697 [link] [comments]
My Databricks Certification exam suspended randomly
Databricks Acquires Row Zero, Bringing Live, Governed Spreadsheets to Genie
submitted by /u/sai-nageshwaran [link] [comments]
Problems with UC MCP connection
We built a supervisor agent where one of the tools is the native databricks MCP connection (which is in beta). Today, the agent doesn’t work. Removing the MCP tool fixes it, so it’s definitely that. The agent response just says “received an invalid and unexpected value from the API: undefined” Anyone else having problems suddenly with the MCP server? I’m able to hit the target MCP URL locally no problem. And even test the UC MCP in databricks successfully. So it’s something about the handshake between MCP and the supervisor agent submitted by /u/pboswell [link] [comments]
Operationalizing Lakeflow Connect:Handling Upstream Schema Evolution & Historical Backfill
Laya off the benchmark: can a zero-shot decision model route real SQL traffic?
Weekend-ish experiment on Databricks. One table, TPC-H orders, 15M rows, living in two places at once: Lakebase (Databricks' managed Postgres), a continuously synced copy with a btree index on the key. Sub-100ms point lookups, useless for a GROUP BY over 15M rows. Delta behind a serverless SQL Warehouse. Great at scans and aggregations, slow at fetching one row. The synced table is the nice part: native continuous Delta to Lakebase Postgres (needs a PK + CDF), so it's one dataset under one Unity Catalog, replication handled by the platform instead of a homemade pipeline. Then I put laya (convaiinnovations/laya, a non-autoregressive zero-shot decision model, vanilla, no fine-tuning) behind a FastAPI Databricks App. It reads each query and picks the engine. MLflow traces input to decision to execution. I submitted by /u/Limp-Park7849 [link] [comments]
Serverless DNS failures after trial upgrade (AWS us-east-2)
From Databricks POCs to Building a Data + AI Community
Lakebase Data API (GCP) returns jwk not found for valid service principal tokens
Databricks Acquires Row Zero, Bringing Live, Governed Spreadsheets to Genie
submitted by /u/orangepunc [link] [comments]
Announcement | Genie One MCP: Give any AI Agent the Right Business Context
Tags no Databricks: Governança, FinOps e MLOps
How to override ABAC policies
Databricks Acquires Row Zero, Bringing Live, Governed Spreadsheets to Genie
Genie App Builder, App Spaces, and Micro Apps released to beta
Yesterday (Sept 23), the app features all of us eager for since they were previewed at Summit were released. In testing, it was not available in all workspaces (cloud and region available differ). Looks like you have to use App Spaces to use these features. App Spaces is a collection of app(s) with some associated boundaries. I'm thinking like if you had a group of apps for a specific team with same users and data and developers, but since brand new don't know if that is exactly how they work. App Builder - A simple UI front end for non-technical users to build an app using Genie. Will be curious to see how this is different than the chat Genie Code interface. Micro Apps - This is the feature I think most of us wanted, the ability to scale to zero, and only pay for actual app consumption with no minimum spend Really excited to see these in the wild. submitted by /u/ecp5 [link] [comments]
Genie Use Cases & easy adaptibility
DQ Management
Hi, we have a lot of data products, how do u design your data quality, and what tools do u use? Idea is that we use dqx, we define rules (ERROR, WARN), and enforce that in pre write of the data, which is to be honest a new thing. Then i would also like to have anomaly checks, where i could compare stats of current batch run to the previous one. How do u design your dq around dqx? And what else do u use? What "flow" do u use? submitted by /u/ptab0211 [link] [comments]
Access to Lakehouse Real-Time (RT) in the preview tab
🚀 Databricks AppQuest Quest 5 – Intelligent Knowledge Base with RAG
Conditionally retrying a task on Lakeflow Jobs
I've entered a new company which uses DAB for all orchestration, and my first task has been to automate certain pipelines that fail due to data unavailability so they retry the fetch later instead of requiring manual runs. Issue is that data fetching errors are one of the many kinds that can be raised during the tasks, and it should be the only retryable one. I'm new to DAB and I've been diging into docs and forums, but I cannot find a native approach or non-patchy workaround to only retry on certain conditions (like based on exit codes). So far, I've come with these candidate solutions: - Just retrying everything (which doesn't makes sense cause if any of the other errors occur, all attempts would fail) - Failing only for the data fetch and propagating fatal errors to the following task, which would check state and fail if errors occurred - like via notebook return values (I don't like this because couples orchestration into business logic, even in a distinct layer of the one that generates the error, plus, the task that would appear as failing would be the successor to the fetching task, which would be shown as successful - not to speak about the hundreds of tasks that would need source code changes) - Failing only for the data fetch and adding an intermediate if/else task that checks if (via notebook return) the previous task propagated and error and stopping the job in such case (this is for now the most likely, as I can keep everything at orchestration level, and it is a bit more explicit even if the fetch task appears as successful, though it still feels like a workaround and tons of jobs would have to implement this new intermediate task) A couple of clarifications: - I cannot implement the retry within the task as fetches would be repeated after some time, and we don't want to have resources running and being billed just for waiting - Everything is done via Lakeflow jobs and with Databricks computer options (even if Spark is not needed) - Most of the fetching occurs by calling external FTP servers, data doesn't directly arrives to our infrastructure like to trigger the pipelines via events Thank you all Edit about explicit behavior: - On data fetch related error, retries occur - On other errors that can come from the task, the whole job fails without retrying anything - On success, next task executes - If failing due to a non retryable error, or due to maximum retry, the job fails without executing downstream tasks submitted by /u/Xandor19 [link] [comments]
Deploying My First Databricks App with Lakebase
Spark Internals on Databricks
Genie agent hyperlink creation
Event driven pipelines with Apache Airflow
For those using Apache Airflow in production, are your pipelines mostly scheduled or event driven If they are event driven, how do you trigger them in practice Do you use Sensors, Airflow Assets/Events, Kafka, EventBridge or something else And for those handling a high volume of events, do you still use Airflow or do you move that part to something like Kafka and Flink submitted by /u/almightysosa888 [link] [comments]
Is Advanced edition required for Serverless pipeline?
Genie Agents citation (bug)
Automated Actions after manually adding Users to Workspace
Hi, in my current project I want to add users manually (no EntraID) to the Workspace and having a pipeline/scripts executed for the users automatically - such as: Creating a catalog based on the user name Granting the user rights to the catalog Loading small amounts of data etc Basically to initiate new user specific actions. I have a few ideas how to realize it, but would like to keep the time from adding the user to execution minimal. Does anyone has experience with such large use cases (in production) ? Thanks for any help! submitted by /u/Prim155 [link] [comments]
How can I configure Lakeflow Connect SQL Server CDC Gateway to use a desired VM type?
zerobus-ingest over kafka
I tested zerobus-ingest over the Kafka-compatible API. It worked fairly well but a little slower than I hoped. I'm guessing that is just the penalty that comes with using Kafka instead of the proprietary SDK. I can't use the proprietary SDK as of now, so I'm working within those constraints. I noticed that this Kafka interface seems to be restricted to json/text, which seems problematic where performance is concerned. Any thoughts on how to scale up and get more throughput? The docs clearly say that batching is the single biggest lever for throughput, but even after playing with batch sizes my data is still not uploading faster than 3 Mbps of json or so. I'm assuming there is no restriction (concurrency/blocking) that would prevent me from running ten kafka producers at a time? Maybe that allows me to get up to 30 Mbps. I'm really tempted to just go back to dropping parquet files in a temp folder. No matter what magic databricks has conjured in zerobus, I find it hard to believe they can compete with the performance of dropping blobs into storage. submitted by /u/SmallAd3697 [link] [comments]
Issues with oidc/v1/token endpoint (REST for zerobus-ingest)
Anyone else work with the zerobus-ingest? There is an auth-flow to get an "oauth token" from a service principal with an "oauth secret". In that flow, you are supposed to submit the authorization details. See docs: https://learn.microsoft.com/en-us/azure/databricks/ingestion/zerobus-kafka I had been requesting access (via authorization_details) to a table as shown in the example: "type": "unity_catalog_privileges", "privileges": ["SELECT"], "object_type": "TABLE", "object_full_path": "main.default.air_quality" ... but it wouldn't work. Was banging my head on it all day. The privileges granted to the service principal for this table were unrestricted " ALL PRIVILEGES ". It turns out that, in addition to "ALL PRIVILEGES", this "token" endpoint wanted me to grant "SELECT " to the service principal in UC governance as well. Else the authorization_details are not valid. This "token" api has a lot of sharp edges. Any chance there is a nuget client to wrap around it and protect us from these unfortunate nuances? Is it documented that ALL PRIVILEGES doesn't actually give us all privileges, if we request "SELECT" by name? submitted by /u/SmallAd3697 [link] [comments]
Databricks AI/BI Get Session User
Back to the Future Part IV: SDP Rewind
Hey everyone! PM on the SDP team and back to share a cool feature that went into Beta recently for Spark Declarative Pipelines: SDP Rewind ! TLDR Think of it like an undo button for your pipelines. You pick a point in time, and it rolls every table, source offset, and operator state in the pipeline back to that instant. Then you fix your code, schemas, or whatever caused the bad data to flow through, restart the pipeline as you normally would, and it replays from the rewind point. This avoids the need to do a full refresh to recover the pipeline state and saves a bunch of $$ and time. Wait, can't I do this already? What's different about this? Sort of. This is best illustrated with a (hopefully) simple example: Auto Loader ingests order files into a bronze table, and a streaming query aggregates them into hourly revenue in gold. At 6 am, a bad batch lands with amounts in cents instead of dollars, so every order since is 100x too big. Today, you can use Delta Time Travel to restore a single table. But your pipeline consists of three things that move forward together: the tables, the Auto Loader's checkpoint (which files it has read), and the gold aggregate's state (its watermark). Tons of places you can mess up: Roll only the tables back to 5 am, but the checkpoint still says every file is read. The pipeline resumes past the bad files and never re-reads them. Reset the checkpoint to re-read those files, but the rows are already in the tables. Now every order counts twice. Fix both, and the watermark on gold is at 9 am. The re-read rows will get dropped. It's easy to see why teams just do a full refresh instead (but this can be very $$). This is exactly what SDP Rewind will help with :) How to get started Move the pipeline to the Preview channel, set pipelines.rewind.enabled to true , and run it once. It generates rewind points automatically, about once an hour, and keeps them for up to 7 days. Trigger a rewind from the pipeline UI . Also check out the companion repo we created that allows you to demo/test this out yourself: https://github.com/databricks-solutions/sdp-rewind-replay Some caveats We currently support only Kafka, Auto Loader, Streaming Tables, and Delta Tables as sources. More coming. Once you rewind, it's on you to restart the pipeline when you are ready. We don't automatically "replay" anything. Would love to hear your feedback on this! submitted by /u/geniebricks_sud [link] [comments]
