Pular para o conteúdo
Comunidade

O que a comunidade está perguntando.

Discussões recentes do r/databricks e da tag [databricks] no Stack Overflow — problemas práticos, dúvidas de integração e casos limite que vale a pena conhecer.

This week

22 questions
Reddit

Solutions Architect Final Round (Build Demo Pitch Round)

Hi all, I have my final round for a Solutions Architect role (the build demo pitch round) coming up this week, and I'm trying to understand what to expect. Are the prompts mainly focused on proposing a data ingestion pipeline / medallion architecture or is the task more complex? Is the focus of the round the code and architecture itself or presentation / delivery skills? Any help would be much appreciated. submitted by /u/mesathink [link] [comments]

00mesathinktoday
Databricks CommunityData Engineering

Guidance Required: Scheduling a Biweekly Databricks Job

00today
Reddit

SpringBoot to Spark

Transitioning from a Spring Boot to a Spark-based application requires a fundamental shift in perspective. With four years of experience in a Spring Boot environment, now entering the realm of distributed computing. I plan to utilize Java with Spark Streaming, and Databricks for job management and deployment. My primary question is how to effectively transition my thought process from Spring Boot to a Spark-based application. Specifically, I am seeking clarity on distinguishing which Java code executes on the driver and which executes on the executor, and if there are any guiding principles to discern this. submitted by /u/Prakhar____Tiwari [link] [comments]

00Prakhar____Tiwaritoday
Reddit

A song about Fabric & Databricks :)

I've spent the last few years building lakehouses on both Fabric and Databricks, and at some point the only sane response was to write a song about it. "The Rift" is Rock & some Latin Rithms about the stuff we all live with: capacity throttling, two catalogs and two governances, mirroring that breaks at night, the CFO counting every CU, and AI agents writing the code we used to write. The punchline in the last chorus is the honest technical truth: whichever side you pick, deep down it's all Parquet. Full disclosure: the music is AI-generated, the lyrics and concept & voice are mine. It's a side project, and I'm planning a whole album about the data world in 2026. Curious which lines land for you, and which side of the rift you're on. Let me know if you all like the song, and please share and follow, more to come :) https://www.youtube.com/@ThePrimaryKeyandtheRedundants submitted by /u/AccordingTale7158 [link] [comments]

00AccordingTale7158today
Databricks CommunityAdministration & Architecture

Lakebase CDF/WAL2Delta cannot write to UC managed storage with Private Endpoin

00today
Reddit

New in Databricks: parse_sql(): SQL into JSON

Parse_sql() function extracts table references, column names, functions, and parameters from SQL. It extracts without running the query. It also reports syntax errors. more news https://databrickster.medium.com/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudektoday
Reddit

Suggestion needed for Panel presentation - Sr. Specialist Solutions Architect (gen ai)

Hi everyone, I've cleared the coding, GenAI, and architecture rounds and have a panel presentation round coming up. HR said they're working on the next steps and putting the task together. Has anyone been through a panel presentation round for a Sr. Specialist Solutions Architect (gen ai) Pre sales? What does the format usually look like, and what kinds of questions should I expect from the panel? Any suggestions please? submitted by /u/Sea-Bid-934 [link] [comments]

00Sea-Bid-934today
Reddit

how are you splitting shared cluster cost between teams?

we have 3 teams (analytics, ds, and a small platform team) all using the same all purpose cluster bc spinning up seperate ones for each was getting messy. now finance wants a monthly number per team and i honestly have no idea how to give them one tags only get me to the cluster level which is useless here. i looked at system tables and theres user info in there but not sure how ppl turn that into an actual $ split thats defensible, especially with the VM side of the bill coming from azure seperately do you just split evenly? by query time? or did you give up and force everyone onto job clusters lol submitted by /u/United_Pay_8651 [link] [comments]

00United_Pay_8651today
HackerNews

Four Databricks Genie Controls That Don't Stop Malicious Skills

10aleiratoday
Reddit

PSA: old Databricks names vs new ones. Which rename got you?

Databricks has renamed a lot of features over the past year, and a lot of tutorials, blog posts, Stack Overflow answers and even AI assistants still use the old names. If you search the docs for the old name, the results can be confusing, and code samples can look unfamiliar even when you know the concept cold. The ones I see trip people up most: Delta Live Tables / DLT → Lakeflow Spark Declarative Pipelines (often shortened to Lakeflow Declarative Pipelines) APPLY CHANGES INTO → AUTO CDC APIs Databricks Asset Bundles / DABs → Declarative Automation Bundles The concepts didn't change much, but the vocabulary did. Two questions for the sub: Has a rename caught you out, at work, in docs, or in a code review? Did I miss any? I'm sure there are more, especially on the ML / GenAI side. I'll keep this list updated with whatever people add. Edit: removed Liquid Clustering from the list. Fair point from a commenter that it's a change in recommendation, not a rename. Still worth knowing though: older material treats partitioning + Z-Order as the default, and current docs recommend Liquid Clustering for new tables. submitted by /u/UnluckySugar610 [link] [comments]

00UnluckySugar610today
Databricks CommunityTechnical Blog

TokenMaxing to ValueMaxing with Databricks

00today
Databricks CommunityAnnouncements

CUSTOMER STORY | PicPay automates fraud triage with Databricks Genie Agents

00today
Databricks CommunityData Engineering

Problems with jobs / GIT repo

00today
Databricks CommunityLearning Events

Virtual Event | The Emerging Blueprint for Agentic Apps

00today
Databricks CommunityCertifications

dhqodh odw qwd

00today
Databricks CommunityCertifications

djqwgiuld q qoiwdh iqw hdoiqw

00today
Databricks CommunityTechnical Blog

Build Your First Databricks App with Codex: A Practical Setup Guide

00today
Databricks CommunityTechnical Blog

How to choose the right AI agent path on Databricks

00today
Reddit

Data Lakehouse with Agentic AIs: A Guide

submitted by /u/codingdecently [link] [comments]

00codingdecentlytoday
Reddit

Data Lakehouse with Agentic AIs: A Guide

submitted by /u/codingdecently [link] [comments]

00codingdecentlytoday
Reddit

Genie is so Dumb and I am tired of Pretending Otherwise

Hey Guys, Data Engineer with 3.2 YOE here, I have been working on and off on databricks for multiple projects across domains like Retail, logistics and Now Pharma. In my current role, I focus on creating end to end pipeline with different data sources which often get refreshed on a bi-weekly basis for pharma related data The client is extremely demanding and has no technical expertise to gauge how much time it really takes to build and maintain something so complex and therefore expect the team to use Claude teams license and also Genie The problem with Genie is that, I can not trust it, It will not listen to my instructions even after having a separate instructions.md which I update on a daily basis, in fact I update this after each session is complete, and no, I do not use AI to update this, every change request that the client and the client's team has goes through my own words of instruction updates, I have an entire markdown file which has contexts for multiple islands of projects, workspace folders and notebooks that I have to maintain and transition to devops team. I also have created 4 different skill family for Genie, in relation to migration, devops, maintainence and review. And it is so disgustingly bad at handling all this context, that it fails me across all 4 areas. I have seen it lie to my face multiple times, Casting columns as null, altering data types without asking consent, never following my work stream and process in a sequential format Having set the wrong notebook paths in a task even though i explicitly tell it exactly what to do. It often confuses streams of works and ends up mixing so many things that remove this knot and fuck up itself costs me so much time!!! It is very poor in adhering to the domain specific instructions that we provide More than 3 requests in a chat, and it is practically useless And branching of chats is the most useless feature that they have introduced. I have created multiple diagnosis queries to check the work of this agent , and it will straight up lie to you, so much, so convincingly, that you know it knows all the biases you have and it even goes a step ahead and just narrates a story that you can provide to your team in the stand-up. My only concern is, after all this bullshit, my team, of nearly 6 (some of them have never worked in databricks or delta table environments like this) was charged $990 dollars in the month of august , for this shit output??? Have some shame Databricks , fix your worthless product or remove that feature or stop charging so much if you are beta testing in actual prod. Today, I am writing this post because of something that I caught live, that pushed me to the edge. Something that is critical, production level issue, which If I was not paying attention while merging could have been a huge disaster, and mind you the pipelines I make are client facing, what is even more distrubing is the fact that it will just make so many unnecessary changes to a simple query or a pyspark function just enough so that the tests are passed. (So basically it does not want to get caught, and makes a mistake so that we can prompt it again to fix this mistake and then Databricks can charge us more, this looks like a dark pattern to me) In your preview (prima-facia), everything is good, the tests are passing, the job is running, but holy!!!, it was casting 3 columns which are essential for downstream processes as null. It boldly suggests that we remove those 3 columns, and or cast them as null. How is this a solution Databricks? The AI is supposed to have more context than me because it is a agent made specifically for Databricks Environment Correct?? That is how it is sold?? And upon spending just 2 mins, I realized that the fix that it is suggesting will blow up critical information that the client team should see in the app because it is not even there in the first place, and all of this is because it can not read the correct notebooks , even if you tag it. It had se […truncated]

00Remote-Grapefruit214today
Reddit

Is Databricks Classic Compute getting too heavy for small workloads?

I’ve been using Databricks Classic Compute for a while, and recently I’ve started wondering whether cluster cold starts are getting noticeably heavier with newer DBR versions. For example, with DBR 18 LTS, I tried a small 2-vCPU VM for a single-node job cluster and hit DriverStartupTimeout after 300 seconds. Databricks even suggests that this commonly happens on instances with fewer than 4 CPU cores. That feels a bit surprising for workloads that are not actually Spark-heavy — e.g. running Python, dbt-core, API calls, or using Databricks mainly as a job runner inside a VNet. I like Classic Compute because of the flexibility and straightforward VNet/private networking. Serverless is attractive for startup time, but in our environment it would mean quite a bit more networking setup. So I’m curious: Have you noticed Classic Compute cold starts getting slower or more resource-hungry across newer DBR versions? Do you now consider 4 vCPUs the practical minimum for a reliable driver? Has anyone benchmarked the same VM size across DBR 14/15/16/17/18? What are you doing for lightweight non-Spark workloads where you still want Classic Compute? Update: Single node DBR 18 LTS + Standard_D4pls_v6 cold start spent almost 11 minutes 'Waiting for resources' ( Azure Japan East ) https://preview.redd.it/ta90ftommmth1.png?width=442&format=png&auto=webp&s=9815bf14c15362cfe526e96e7904557da11bb839 submitted by /u/bobjia-in-tokyo [link] [comments]

00bobjia-in-tokyoyesterday

Last week

116 questions
Databricks CommunityData Engineering

I built an open-source check that stops AI agents from inventing column names in SQL.

00yesterday
Reddit

Databricks ai_decide explained in 5 minutes!

submitted by /u/ConstantNo2668 [link] [comments]

00ConstantNo2668yesterday
Reddit

App Spaces: governance for apps

Workspace admins, thanks to App Spaces, can define who can create and use apps in a given space and also set policies for apps, for example, which API scopes can be used. more news https://medium.com/databrickscommunity/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudekyesterday
Databricks CommunityCertifications

Exams with additional fees are included?

00yesterday
Reddit

Where do i start if i want to learn databricks concepts. I want to get start using my learnings to apply at work.

submitted by /u/Time_Meringue_1300 [link] [comments]

00Time_Meringue_1300yesterday
Databricks CommunityMVP Articles

Ingest data from Microsoft Outlook

00yesterday
Reddit

Databricks ai_decide: The AI Judge That Never Writes a Word

Use Databricks ai_decide() to grade every ai_query output in SQL and decide what gets sent, without a second LLM submitted by /u/Lenkz [link] [comments]

00Lenkzyesterday
Databricks CommunityDatabricks Free Edition Help

My data saved in the account lost

00yesterday
Databricks CommunityData Engineering

Databricks SDP (Spark declarative Pipelins) Overwrite table

00yesterday
Reddit

Millions in Vacuumable files, but only 10-20 GB recovery size

Hello Folks. Wanted your opinion on something - I am seeing numerous cases of our tables having millions of redundant vacuumable files, but total recoverabke capacity is only 10-20 GB . Is it worth Vacuuming them? My sense is that the DBU and ADLS listing charges will be more than cost of that storage. Is there any instance where Vacuuming will pay off. Also, if deciding to Vacuum - what is better, Job Compute or Serverless SQL Small Size. Thanks. submitted by /u/Lumiaman88 [link] [comments]

00Lumiaman88yesterday
Databricks CommunityCommunity Articles

Built a live Helsinki tram tracker on Databricks with Lakeflow, Lakebase, Genie and Apps

00yesterday
Databricks CommunityCertifications

Certificate Not Received After Passing Databricks Certification Exam

00yesterday
Reddit

Cross-Catalog Sync: Iceberg on Polaris, Glue, and Unity

submitted by /u/codingdecently [link] [comments]

00codingdecentlyyesterday
Reddit

Building A Data App with dashboard and agentic capabilities in dbx

Hi all, I’m currently working on a POC and im given the flexibility by my boss to explore using any tool (im considering mainly databricks or power apps now) to build out a “dynamic” data app to replace powerbi . I am fully aware that dbx ai/bi dash does not has nearly 100% capabilities of that the power bi dash but I’m just trying to showcase a different experience of reporting that leverages AI and not POWERBI (boring! to leaderships). For more context, im working with SAM/CMDB EoSL data that is all sitting in dbx unity catalog already. My expectations is something simple that looks like a dashboard but i can utilize genie to communicate pr generate any sql findings within the dashboard context. Something easy to initiate and maintain in the future. Additionally, it would also be best that the views created that caters to this app can be reused as native queries to powerbi in case the need in the future that id still have to migrate it there. End product expectations: The end product is expected to be easily accessible without any additional access required for any user that wants to access the app or genie. Think of leaderships and executives using the dash mainly. It should be something that is easy to access and managed within a workspace like what powerbi serves currently which is what all the stakeholders are currently used to already. I am no expert in Databricks but have been an active user for about a year now. All the best to everyone, cant wait to hear what you guys have in mind in the comments! (this is my first reddit post btw) submitted by /u/Junior-Confusion-691 [link] [comments]

00Junior-Confusion-691yesterday
Databricks CommunityGenerative AI

How are you attributing agent cost per successful task, not just per token?

002d ago
Databricks CommunityGet Started Discussions

[[Guía@Delta-México]] PAso a PAso ¿Cómo hablar con Delta en México?

002d ago
Reddit

Let AI Decide: SQL Decision Making with Databricks ai_decide()

Databricks just introduced ai_decide(), a new AI function now available in Beta. Give it text and your questions, and it returns a category, a probability, or a score directly in SQL. Use those results to drive decisions in your workflows: route requests, prioritize work, or determine when human review is needed. submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek2d ago
Databricks CommunityGet Started Discussions

PowerBI refresh disabled

002d ago
Databricks CommunityTraining offerings

Exams with additional fees are included?

002d ago
Reddit

Omnigent 0.16.0 + Hermes on macOS: "hermes did not accept the message" and .env/auth.json copied into every session. Anyone fixed this?

I'm new to this, so apologies if I'm missing something obvious. Setup: macOS (Apple Silicon), Omnigent 0.16.0 installed with uv tool install omnigent (Python 3.12), Node 22, tmux 3.7, Hermes Agent installed via git (up to date). Running fully local on 127.0.0.1:6767. Problem 1: first message gets dropped. When I start a new Hermes session from the Omnigent desktop app, it fails with: inner executor error: hermes did not accept the message (the TUI may still be initializing); no new transcript row appeared after two delivery attempts then Required terminal exited unexpectedly; the session runtime is no longer available. The logs show Hermes is still starting up (installing dependencies, browser tool checks) when Omnigent pastes the message, and then Hermes exits with a KeyboardInterrupt. Running omnigent hermes in one-shot mode from the terminal works fine. Problem 2: credentials copied per session. Every Hermes session creates a new folder under the macOS temp dir ( .../T/omnigent-501/hermes-native/ /hermes_home/ ) with a copy of my Hermes .env and auth.json , plus about 1–2 GB of runtime. After one afternoon I had 19 folders (22 GB) and about 30 copies of my keys sitting in temp. Questions: Has anyone gotten the Omnigent desktop app to reliably start Hermes sessions? Is there a setting to give Hermes more startup time? Is there a way to make Omnigent use an existing Hermes profile without copying its .env / auth.json each session (for example, a shared HERMES_HOME or a secrets manager)? Is either issue fixed in a newer version, or is there a GitHub issue I should follow? I'd like to eventually run a dedicated Hermes profile (one with sensitive API keys) through Omnigent, so the credential copying is the big blocker for me. Thanks! submitted by /u/Kwontum7 [link] [comments]

00Kwontum72d ago
Databricks CommunityCertifications

Rescheduled my exam of My Data Brick ML Engineer Associate Exam

002d ago
Databricks CommunityGenerative AI

Premium PAYG account stuck at TRIAL_VERIFIED (rate limit of 0 on resold models) — how do I get promo

002d ago
Reddit

How do you guys handle schema changes in Databricks when the source keeps changing?

Like the source suddenly adds a new column, changes a column type, or removes something. Do you make the pipeline handle these changes automatically, or do you just let it fail and fix it? Iam.. curious what approach actually works better in real projects, especially when the source keeps changing. submitted by /u/Bhanuprakash_1947 [link] [comments]

00Bhanuprakash_19472d ago
Databricks CommunityData Engineering

Analyze

002d ago
Databricks CommunityData Engineeringanswered

Lakeflow connect

002d ago
Databricks CommunityGet Started Discussions

Noxivam – En enkel kapselrutin för vuxnas dagliga välbefinnande

002d ago
Databricks CommunityAdministration & Architectureanswered

Serverless compute is not available in my account/workspaces

003d ago
HackerNews

Postgres is now the most efficient search database at scale

30moonikakiss3d ago
Reddit

How to fit ~1GB+ embedding model into a 2Gi Kubernetes pod? Getting OOMKilled

Hi, Deploying a FastAPI to Kubernetes that uses a multilingual sentence-transformers embedding model (ONNX backend, CPU only). My source data lives in a Delta table in Databricks, and the app reads from it and generates embeddings. The app takes user input (text) at embeds it, and compares it against stored embeddings for semantic similarity. So at least the query embedding has to happen live. Pod resources: - CPU: 1 request / 2 limit - Memory: 2Gi request = 2Gi limit (platform policy requires memory request and limit to be 1:1) Current docker setup: - Multi-stage Docker build (python:3.11-slim) - CPU-only PyTorch - The model is downloaded at build time, so it's baked into the image Image breakdown: - HF model cache: ~1.1 GB - torch: ~650 MB - pyarrow, scipy, transformers, pandas: ~100–150 MB each - Plus sklearn, onnxruntime, mlflow, and others The pod gets OOMKilled at startup or shortly after. With a ~1GB model, torch, pandas/pyarrow, and the ONNX runtime session all in one process, I think I'm just over 2Gi. What I'm trying to figure out on where do you store large models in production? Would love to hear what setups have worked for you. Thanks! submitted by /u/runningnozone [link] [comments]

00runningnozone3d ago
Reddit

Feature Request: pipeline_task should have "no wait" option

Sometimes we want to trigger a pipeline (such as database sync) without having to wait for its completion. We can still subscribe to notifications on the pipeline itself to react to an unexpected error, so the pattern should be fine. submitted by /u/CarelessApplication2 [link] [comments]

00CarelessApplication23d ago
Databricks CommunityTraining offerings

The trendtrack promo code is BESTCOUPON20

003d ago
Databricks CommunityAdministration & Architecture

Resource limits

004d ago
Reddit

I am trying to build an interactive dashboard on the underlying Databricks. Which of these are the best?

I’m thinking of 4 options here 1. Build an MCP (a custom MCP) that can access custom tools on Databricks and interface it on Claude.ai or Claude desktop Advantage- Claude is very good at inferencing, multi-turn conversation and multi-step processing Disadvantage - custom MCP and tools needs to be built accurately and validated. It should have full context of schema and unity catalog Build a semi-custom MCP - this will use “askGenie “ as one of its tools with additional custom tools Advantage- complexity decreases as we leverage genie space Disadvantage- double inference by genie and Claude Use custom Databricks connector in Claude. Not sure if this uses genie and therefore double inference but it’s more reliable than custom build because this is a native offering by vendor Use only genie and build custom dashboard without needing Claude interface What’s the thought on this? submitted by /u/Bala_Devaraj [link] [comments]

00Bala_Devaraj4d ago
Reddit

databricks TaskValue and Lakeflow

Databricks taskValues can use a Python list as input into a For each loop in Lakeflow job. You can generate the list dynamically and run the same task for every country, table, file, etc. more newshttps://databrickster.medium.com/databricks-news-serverless-genie-code-ltap-lakeflow-61853d8e422a submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek4d ago
Databricks CommunityGet Started Discussions

PARTITIONED BY, Liquid Clustering, and OPTIMIZE: when should you use each?

004d ago
Databricks CommunityAnnouncements

Join the Live AI/BI + Genie Launchpad (Oct 6–8)

004d ago
Databricks CommunityAnnouncements

CUSTOMER STORY | BookMyShow scales self-service analytics with Genie

004d ago
Reddit

The Breakdown: Databricks

submitted by /u/Roadtochessmaster [link] [comments]

00Roadtochessmaster4d ago
Reddit

'DATABRICKS DEN' - NEW COMMUNITY

Over the past few weeks, I've been building something special. Every day I speak with Databricks experts across Europe. After so many of those conversations, the obvious question was: why not build a community around it? So today I'm launching 'Databricks Den'. Whether you're looking for your next contract or just want to stay close to what's happening in the Databricks market, the Den is for you. Join now to catch the first round of insights: 🔗 Link to the Den Looking forward to seeing you there 👊 submitted by /u/Reuben_UMATR [link] [comments]

00Reuben_UMATR4d ago
Databricks CommunityGenerative AI

Gemini Foundation Model endpoint unavailable in 14-day commercial trial

004d ago
Databricks CommunityGenerative AI

Routing between GLM 5.3 Flash and GLM 5.3 with ai_decide: a cheap router in front of ai_query

004d ago
Databricks CommunityMachine Learning

Document AI is a pipeline, not a model

004d ago
Reddit

Setting my Expectations for Microsoft CSS support in Azure Databricks (2026)

I recently moved back to using Azure Databricks after a multi-year hiatus. And I opened my first support ticket this week. My prior support experiences were long ago (in 2021). I don't remember these support experiences being particularly terrible. But nowadays there are some Microsoft platforms where their support work is subcontracted to remote organizations like Mindtree-India. These organizations (one or more of them) may provide the first, and second tiers of support before any FTE from Microsoft takes the reins. Is that how things will work in the context of Azure Databricks? If/when the FTE at Microsoft is unable to help with a ticket, how long will it take for the support to be transferred from Azure Databricks to Databricks? I have a sev B ticket on "standard" support at the moment and I thought a human would engage within a few hours but that isn't happening. I will probably just give up for today. Can anyone offer tips about how to navigate support tickets in Azure? Any advice would be greatly appreciated, and I think it could make the difference between a ticket that lasts two days or two weeks. Thanks. submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36975d ago
Databricks CommunityAdministration & Architecture

Budget override option not available

005d ago
Reddit

The Four Architectures That Make AI Work

submitted by /u/Berserk_l_ [link] [comments]

00Berserk_l_5d ago
Databricks CommunityCommunity Articles

Row-level security for a RAG agent: what Unity Catalog enforces, and what you have to build

005d ago
Databricks CommunityData Engineering

Lakeflow connect SQL server ingestion can be made elastic?

005d ago
Databricks CommunityMachine Learning

MLflow traces accepted by StartTraceV3 (200 OK) but never stored: GetTrace returns NOT_FOUND

005d ago
Reddit

system.query.history is now Generally Available in Databricks

Until now, answering "which query is eating our warehouse?" or "who changed that table?" meant clicking through the Query History UI one workspace at a time. Now it is a SELECT. The table logs every statement run on SQL warehouses, serverless compute, and Lakeflow pipelines, across all workspaces in the same region. Read more on LinkedIn: https://www.linkedin.com/posts/cenh_databricks-dataengineering-systemtables-ugcPost-7511094425977016320-L7uj/?utm_source=share&utm_medium=member_desktop&rcm=ACoAABmJHrsBNAC3x3H1M58JRKoHv_l4D61n0-8 submitted by /u/Lenkz [link] [comments]

00Lenkz5d ago
Databricks CommunityWarehousing & Analytics

How to build FinOps chargeback by cost centre — allocating the whole cloud bill, not just your DBUs

005d ago
Databricks CommunityCommunity Articles

The 32nd Column: Why Your Z-Order or Clustering Key May Be Doing Nothing

005d ago
Reddit

Databricks Lakebase Search - GA Announcements

submitted by /u/Wise_Ear_4064 [link] [comments]

00Wise_Ear_40645d ago
Databricks CommunityAnnouncements

Announcement | Five ways marketers can use Genie One

005d ago
Databricks CommunityData Engineeringanswered

At what data size do you stop reaching for Spark?

005d ago
Databricks CommunityData Engineeringanswered

Photon enabled but a large share of the plan is falling back, cost up and runtime flat

005d ago
Databricks CommunityData Engineeringanswered

Streaming read fails with "Detected a data update" after a restatement job touches the source

005d ago
Databricks CommunityWarehousing & Analytics

Unable to connect to Tableau Cloud

005d ago
Reddit

Serverless - how is this OK for a "fully managed" product?

Notebook that's run fine for months suddenly started failing with 'No Module Found 'chardet' Nothing changed on our end. Seems that serverless environment v6 dropped chardet from the base requirements, and since the notebook wasn't pinned, the "default" just rolled forward underneath it. No warning, and the release notes don't even mention the removal. I get the argument of pin your environment, declare your dependencies. But serverless is sold as the "no config, we manage it for you" option. Silently removing packages from the default environment mid-week, with zero notice, feels like a breaking change shipped as a patch? submitted by /u/OneSeaworthiness8294 [link] [comments]

00OneSeaworthiness82945d ago
Reddit

Lakeflow Jobs update! Task output in foreach now available in Beta!

If you run a ForEach task in Jobs each iteration now sets a task value and a downstream task can read all of them back as on ordered array. This feature just landed in Beta - try it out! Before this, iteration outputs weren't accessible to downstream tasks. If you wanted a later task to see per-iteration results (row counts, status and output paths), you had to write each result to a table and read it back and maintain that yourself. You also lost the built-in observability Jobs gives normal task values. How it works: Inside the iteration, set a value like you always would: dbutils.jobs.taskValues.set(key="result", value=row_count) In a downstream task, read the nested task's key. You get one array across every iteration: results = dbutils.jobs.taskValues.get(taskKey="process", key="result") # results == [1, 2, 3] Or reference it as a parameter: {{tasks.process.values.result}} An iteration that never set the key keeps its slot as null: # [1, None, 3] https://i.redd.it/clkd42sm3nsh1.gif Worth knowing: Beta, needs Databricks Runtime 15.4 LTS or above Python notebooks only The assembled array caps at about 48 KB (49,344 characters) per parameter value A single parameter value can hold up to 3 aggregated references Docs: https://docs.databricks.com/aws/en/jobs/task-values#read-values-from-a-for-each-tasks-iterations And yes we are working on increasing parameter value limits! Please let us know if you have any feedback! submitted by /u/saad-the-engineer [link] [comments]

00saad-the-engineer5d ago
Databricks CommunityAnnouncements

CUSTOMER STORY | ANA Seguros turns insurance data into faster decisions with AI/BI and Genie

005d ago
Databricks CommunityCertificationsanswered

Quick clarification about GENAI or Context engineering certification

005d ago
Databricks CommunityTechnical Blog

Step-by-Step: Building a Vacation Rental Operations App with AppKit

005d ago
Reddit

Azure Databricks, two-week-old account: partner pay-per-token endpoints (Claude, Grok) have never once succeeded — "Databricks-set rate limit of 0" — and the Assistant has "no daily token allowance". Open models work. What gates this?

Azure Databricks, Premium, one account with three workspaces, Unity Catalog, Azure Marketplace billing (no card, no trial credits left). The account was bootstrapped on a trial workspace about two weeks ago and moved to Premium six days later. I've spent two days on this and want a sanity check from anyone who has seen it. The facts, from system.serving.endpoint_usage: • Every open-weight pay-per-token endpoint has worked since the trial: gpt-oss-120b has ~50 successful calls going back to the trial period, Llama 3.3 70B likewise, zero refusals. • No partner pay-per-token endpoint has ever returned a success on this account. The very first call anyone made to databricks-claude-sonnet-5 was refused, and every call since, on every Claude endpoint and on Grok, in all three workspaces, for users and service principals alike: 403 PERMISSION_DENIED: The endpoint is temporarily disabled due to a Databricks-set rate limit of 0. • It is not our AI Gateway config: I removed every rate limit from one Claude endpoint and invoked it — same 403. The endpoints' config shows nothing else. • The notebook Assistant worked during the trial. On the paid account it refuses everyone, admins included: *"This workspace has no daily token allowance for the assistant."* Underneath, /ajax-api/2.0/conversation/llmproxy/ returns 429 {"type":"daily_token_limit_reached","daily_limit_tokens":0,"message":"…Daily token limit of 0 reached. Resets at midnight UTC. (DTB)"} and it recomputes to 0 every midnight. • Once, the Databricks Apps in the workspace were stopped by the platform with *"App compute was stopped due to workspace or account status"* while every workspace showed RUNNING. apps start brought them back. Red herring, for completeness: we also created a Unity AI Gateway per-user budget whose first version had a $0 threshold with BLOCK_USAGE. That blocked Genie for a day, exactly as documented, and was fixed. It was created a day *after* the first Claude refusal, so it isn't the cause of any of the above. What I've found: two Community threads describe "Databricks-set rate limit of 0" as a workspace trust-tier gate — trial-born accounts sit in TRIAL_VERIFIED, pay-per-token partner models are gated to PAYABLE_VERIFIED, a card alone doesn't flip it, only Databricks Sales/Support moves the tier. Both were AWS/personal accounts, neither resolved on-page. The account console's own API reports our account's feature_tier as STANDARD_W_SEC_TIER although the workspaces are Premium; I can't find what that field means. Questions • Does Azure Databricks with Marketplace billing go through the same PAYABLE_VERIFIED gate? Does it clear on its own after the first settled invoice, or does someone have to move it? • If you were moved: who did you contact (account team, or "Contact us" in the console) and how long did it take? • Is the Assistant's "daily token allowance" the same gate, or a trial allowance that simply goes to 0 when the trial ends on an unverified account? • Does feature_tier: STANDARD_W_SEC_TIER on the account object mean anything to anyone? submitted by /u/sumit671 [link] [comments]

00sumit6715d ago
Databricks CommunityData Engineering

Built my first Databricks App with Lakebase

005d ago
Databricks CommunityWarehousing & Analytics

Dashboard DAB Deployment: default_catalog and default_schema do not work for metric views anymore

005d ago
Reddit

Let's talk about AI rate limits

Databricks Unity Gateway is great and getting even better very quickly. It now has nearly everything we could hope for. New frontier models are usually available within 24 hours of release. But...the models aren't really available . They all come with extremely low rate limits, way too low to support a team of engineers in any sizable enterprise using the models. If you have a sizable agents.md even 1 person hits the limits. Why is it like this? Sure we can raise a ticket to request higher limits but that takes weeks and the highest we've gotten is 1M tok/min. We can (and do) go to Bedrock and get 5M+ tok/min by default, no special request needed. Is the expectation that we hook up other providers? Are the key frontier labs (OpenAI, Anthropic, Google) limiting Databricks as a whole? For such a great product this seems like a glaring limitation that's holding us back. Make it make sense. submitted by /u/degenbets [link] [comments]

00degenbets6d ago
Reddit

DB as a SIEM

My organization is considering replacing our current SIEM with DB. Wondering of anyone has done it and looking for feedback. Specifically looking for any gotchas, architecture recommendations. Cheers! submitted by /u/Environmental_Arm370 [link] [comments]

00Environmental_Arm3706d ago
Reddit

Spark and observability

submitted by /u/ZenithR9 [link] [comments]

00ZenithR96d ago
Reddit

Writing to externally managed Iceberg tables from Databricks: am I missing something?

I recently joined a new company as a data architect and I'm trying to understand how things work here today (new to Databricks, main experience with RedShift and BigQuery), and I'd like corrections from people who use Databricks daily. I've been testing how Databricks handles Iceberg tables that are managed by Glue. What I'm seeing: Can read tables as Databricks' federation lets me mount the external catalog and query the tables as foreign tables. Can't write (is this accurate?) INSERT, MERGE, and UPDATE operations against those foreign Iceberg tables fail, and the Databricks docs seem to confirm this. Question: Has anyone gotten writes to non-UC Iceberg tables working from Databricks? That could be through a config, a preview feature, or an open-source Spark Iceberg runtime with a REST catalog on a cluster. If I've got any of this wrong, please correct me as I am still getting up to speed with the platform. submitted by /u/DonutSignificant4174 [link] [comments]

00DonutSignificant41746d ago
Databricks CommunityCertifications

DevOps Fundamentals: Lecture - DevOps Fundamentals: Repeated word

006d ago
Databricks CommunityData Engineering

1.Not able to access interent (gcp databricks) 2.both ingress and egress access given

006d ago
Databricks CommunityData Engineering

Tip: Why your Delta MERGE fails with "multiple source rows matched" (and the simple fix)

006d ago
Databricks CommunityData Engineering

Azure to AWS

006d ago
HackerNews

The lakehouse serving fight is on

50dashdoesdata6d ago
Reddit

Serverless Interactive and Realtime Inferencing costs scaling rapidly

Hello folks, I have recently started a new role as Data Platforms leader, and while going through our monthly Databricks billings - two things stood out, both mentioned in Title, costs for them scaling rapidly month on month. Serverless Interactive compute, by default has access for Admins inherited, and people now dont want to go back to the slow time taking All Purpose Compute starts. Serverless Realtime Inferencing has me confused, as I cannot see any such compute - is that cost for using Genie? Is yes, can Genie access be turned off for everyone in the Workspace? Just these two items reduced will save 20-25% costs for us. Thanks in advance for any pointers. submitted by /u/Lumiaman88 [link] [comments]

00Lumiaman886d ago
Reddit

Can a Delta table be “correct” even when the source and target row counts match?

Suppose the source and target both contain 100M rows. The row counts match, but the target could still have: Missing records Duplicate records Incorrectly transformed values Changed or truncated values For a large Delta table, what checks would you use beyond row counts to prove that the source and target data actually match? Would you use hashes, key-level reconciliation, aggregates, sampling, or something else? submitted by /u/No_Ambition8323 [link] [comments]

00No_Ambition83236d ago
Databricks CommunityData Engineering

Implementing a "Zero-Bus" Architecture with Unity Catalog

006d ago
Databricks CommunityData Engineering

Implementing a "Zero-Bus" Architecture in a Lakehouse: Best Practices for Shared Dimensions?

006d ago
Databricks CommunityAnnouncements

Announcement | The Genie One MCP is now Generally Available

006d ago
Databricks CommunityGenie Hub

How to make each person see only their specific data (i.e. their own rows in Genie) with RBAC & ABAC

006d ago
Reddit

Tag automation with databricks

No table description? Not queried for 90 days? See how to automatically tag the table and notify its owner in my new article. https://medium.com/databrickscommunity/clean-up-the-mess-with-tag-automation-72ae119b6f35 https://www.sunnydata.ai/blog/databricks-unity-catalog-automated-tagging submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek6d ago
Databricks CommunityCommunity Articles

Solution Accelerator Series | Incident Investigation Using Graphistry

006d ago
Databricks CommunityGet Started Discussions

Lessons learned from configuring Genie Agents

006d ago
Databricks CommunityCertifications

Certification Support Request #01015865 – No Response Since September 14

006d ago
Reddit

Data Quality in Microsoft Fabric – Native features vs. external tools (like Databricks Expectations)?

Hey everyone! I’ve been working a lot with data quality frameworks in Databricks (leveraging things like Delta Live Tables expectations or custom validation notebooks), and I'm curious about how the community is handling Data Quality in Microsoft Fabric. Does Fabric have a native, built-in feature for data quality/expectations (similar to what Databricks offers), or are most of you relying on external tools, libraries (like Great Expectations/Soda), or custom PySpark/SQL validation notebooks inside Data Engineering pipelines? How are you currently enforcing quality checks before data hits your Gold/Semantic layers in Fabric? Would love to hear your approaches and best practices! I read this documentation here: https://learn.microsoft.com/en-us/purview/unified-catalog-data-quality-fabric-lakehouse?WT.mc_id=510336 , but it doesn't say much. submitted by /u/SuperbNews2050 [link] [comments]

00SuperbNews20506d ago
Databricks CommunityGenerative AI

WORKSPACE HAS NO DAILY TOKEN ALLOWANCE

006d ago
Databricks CommunityAdministration & Architecture

Deployment Jobs structured yaml.

006d ago
Reddit

Time to Swap the Cookies for Jetfuel - New Dataset & New Databricks Genie Tutorial

Most of you know samples.bakehouse . Great for a first query. Perfect for a quick demo. But after years of cookie sales, it's a little overbaked. Time to swap the cookies for jet fuel. ✈️ Together with the OpenSky Network , I brought a full day of global air traffic to Databricks Marketplace: 696 million real ADS-B position reports, messy just like real life. Myself, I used Genie for the whole journey: EDA, data exploration, a Apache Spark Declarative Pipeline, and a Lakeflow Job. Then I went a step further and read the same Marketplace data with open-source tools only, using OpenSharing and pandas. The result is this hands-on tutorial: Marketplace + Unity Catalog: get the data as a governed table Genie Agents: find anomalies in plain English Genie Agents: explore and visualize with maps and charts Genie Code: a Spark Declarative Pipeline, bronze to gold, with data quality rules Genie Code: a Lakeflow Job with schedule, retries, and email alerts Databricks Apps: your coding agent, governed by Unity Catalog OpenSharing: the open-source client in VS Code with pandas Everything runs on Databricks Free Edition (free, no credit card). 📖 Tutorial: Databricks Genie for Data Engineers and Data Scientists 💻 GitHub: databricks/tmm/DSDE-Genie-Tutorial 🛫 Dataset: OpenSky Network full-day dataset on Marketplace What's the first thing you'd query in a day of global air traffic? P.S. For the record, we still love bakehouse! 🍪❤️ [Disclaimer: I'm one of the two people who baked it.] submitted by /u/CompetitiveBet8978 [link] [comments]

00CompetitiveBet89786d ago
Databricks CommunityGet Started Discussionsanswered

How to reduce query latency?

006d ago
Reddit

How to validate delta tables in Databricks?

im working with delta tables in databricks and wanted to know how ppl usually validate the data after an etl job runs. apart from checking if the job completed, do we have to check things like record counts, duplicate, nulls etc Like if the source has around 100k records, how do you make sure the expected records reached the Delta table? for incremental loads, do you also check that only the expected records were added or updated? And how do you handle schema changes, like a new column being added or a data type changing? Do you build these checks directly in databricks itself? Would be interested to know what checks ppl actually use in production. submitted by /u/Klutzy_Solid5200 [link] [comments]

00Klutzy_Solid52006d ago
Databricks CommunityData Engineering

How to bring the project id info into the databricks billing usage table on GCP

006d ago
Reddit

Avoid Unity Catalog external location permission bugs by using managed storage

Hey, I work as a Databricks Engineer at Abilytics, so here's my take. If you are migrating existing tables to Unity Catalog, defining external locations incorrectly will break your grants. When you register an external location via CREATE EXTERNAL LOCATION, Unity Catalog validates your storage credential against the underlying cloud storage container. However, users often forget that read and write access on the external location is distinct from the IAM role or service principal trust policy. If a user has SELECT on a table, but lacks BROWSE on the external location, queries against parquet files directly via path-based access will fail with permission denied, even if the table-level grant succeeded. Always verify your storage credentials have explicit container-level policies attached before mapping external volumes. Furthermore, remember that path-based access like spark.read.load('s3://my-bucket/path') requires explicit external location grants, whereas managed tables abstract this entirely. Prefer managed tables default storage roots to bypass manual external location ACL synchronization issues entirely. Happy to go further — more on Databricks | Platform Engineering | AI at Abilytics , AI Systems, Data Engineering & Platform Engineering Services. submitted by /u/AbilyticsEng [link] [comments]

00AbilyticsEng6d ago
Reddit

Stop using standard worker nodes for intermittent batch jobs

Hello, I'm part of the Databricks Engineer team. If you are running short-lived batch pipelines on standard clusters, you are wasting money on idle instance termination times. Switch your job clusters to use spot instances by configuring the instance pool or setting the spot bid policy directly within the Databricks Jobs API or UI cluster settings. When configuring a job cluster, set the spot instance policy to use spot with fallback to on-demand to prevent job failures when spot capacity is scarce. For intermittent batch jobs and maintenance tasks like OPTIMIZE and VACUUM, using spot instances significantly reduces compute costs compared to persistent multi-node on-demand clusters. Ensure your retry logic handles node preemptions gracefully, and leverage job clusters rather than all-purpose clusters so the cluster terminates immediately after the batch execution completes, eliminating idle billing. If this is useful, there's more on Databricks | Platform Engineering | AI at Abilytics — AI Systems, Data Engineering & Platform Engineering Services. submitted by /u/AbilyticsEng [link] [comments]

00AbilyticsEng6d ago
Databricks CommunityCertifications

Courses still showing as in progress

001w ago
Reddit

Data fixes on a databricks schema. Thoughts?

Hi all, I have a databricks schema of a few hundred delta tables that need some data fixes for specific records in each of those tables . This schema itself is raw data and gets ingested into some downstream data tables and the fixes have been requested by business. Now, importantly, this schema is the most upstream source for this data - the SQL db it was ingested from no longer exists. I had thought about doing the fixes via a transformation layer in the pipelines that ingest this raw data but given how many tables there are with records that need updating I don't really want to create hundreds of new 'data fixed' tables. I also read that it's best to fix something as upstream as possible. Given these are delta tables, rolling back should theoretically be possible if something goes wrong. Anyway, that's my rationale for making changes to the source tables. My question is what your preferred method is to make data fixes? I obviously need something where its easy to rollback if needed. I can obviously just achieve this with python migration scripts and use delta timetravel in case something goes wrong, but wonder if there are recommended libraries or tools for the job that have what I need out of the box? Should I also keep an unchanged copy of this schema or is that too redundant? submitted by /u/Spooked_DE [link] [comments]

00Spooked_DE1w ago
Reddit

Databricks Micro Apps, App Spaces and Genie App Generator

Databricks launched App Spaces and Serverless Micro Apps at the end of last week (in Beta). Over the weekend, I tested migrating two of my production Databricks Apps to App Spaces. App Migration 1: Blocked by Zero Egress My Data Portfolio Project Creator required internet access to pull data stacks from live job postings. Then, the LLM needs internet access to research for open data sources to use. In standard Databricks Apps, this runs cleanly. In App Spaces, there is zero external internet egress including for LLMs. The app can reach internal workspace resources, but it cannot touch the outside web. If your app relies on third-party APIs, external databases, or web scraping, App Spaces is a non-starter until Databricks opens network egress. Workload 2: Success on Internal FinOps My client-facing DBU cost observability app reads workspace usage data and writes directly to Lakebase. Because it requires zero external network calls, the migration worked. Both the micro app compute and Lakebase scale to zero when idle. Cold starts take roughly 30 seconds (though in beta, you occasionally need a quick browser refresh once it spins up). For internal, low-frequency administrative tools, this turns a continuous monthly compute bill into pennies. (+ App Spaces, Genie App Generator, and Micro Apps are free while in beta) The next part is less about App Spaces and more just general best practice for Databricks Apps that I see people miss. Stop Using Delta Lake as an OLTP Database Databricks Apps are software applications, not batch analytics notebooks. If your app writes application state, session data, or row-level CRUD directly into analytical Delta tables, you need to rethink that design. For app transactions, use Lakebase (serverless Postgres which also scales to 0). Both are governed under UC, but Lakebase gives your app the low-latency transactional engine that application engineering actually requires. My Verdict on the Beta App Spaces solves the idle compute problem that has plagued Databricks Apps since launch. But until Databricks allows us to deploy directly from existing Git repos and opens external network egress, it remains limited to internal-only use cases. Curious on other peoples experience... How has it been for others? submitted by /u/OkImprovement7010 [link] [comments]

00OkImprovement70101w ago
Reddit

Business data context is the key to better answers.

submitted by /u/ConstantNo2668 [link] [comments]

00ConstantNo26681w ago
Databricks CommunityData Engineering

Declarative Pipeline Expectations - Quarantining Bad Records

001w ago
Reddit

No compute fleets in Azure Databricks?

I see that AWS has supported spot-priced compute fleets for a few years. But Azure still doesn't use fleets to select spot-priced VM's (yet). Customers are forced to set up our own spot-priced pools and manage a fleet of them independently. Two years ago someone named Walter (an FTE at databricks?) said "currently on development but no ETA " https://community.databricks.com/t5/administration-architecture/compute-fleets-on-azure-databricks/td-p/104754 So what is the reason for so much delay in Azure? Is there some sort of conspiracy going on where Microsoft and Databricks are teaming up to give us only expensive compute (like serverless) and restricting our ability to access other alternatives? This is getting a bit frustrating. submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36971w ago
Databricks CommunityGenerative AI

Salesforce Hosted MCP + UC HTTP connection — stops working after ~1 hour?

001w ago
Reddit

Workshop on bringing systematic evaluation and MLflow tracking to LLM apps, Oct 3

If you're already using MLflow for experiment tracking on the ML side, there's a decent chance your LLM work hasn't caught up to the same standard yet. Most teams are still hand-tuning prompts and eyeballing outputs while everything else in the stack gets versioned and tracked properly. Serj Smorodinsky and Brett Kennedy, co-authors of a book on LLM applications, are running a live 3-hour workshop on Oct 3 that applies that same rigor to LLM development. What's covered: Programming LLM behavior with DSPy signatures and modules instead of hand-written prompt strings Building a baseline classifier live, from a real task Constructing an evaluation dataset with task-specific metrics so "better" is measurable, not a feeling Reading failure patterns directly out of the eval results Few-shot and instruction-level optimization applied systematically on top of the baseline Experiment tracking and LLM trace management through MLflow, so every run is reproducible and comparable Saving and reusing optimized DSPy programs across projects Communicating LLM reliability to stakeholders, which usually gets skipped entirely It's basically the MLflow discipline this community already applies to models, extended to prompts and LLM pipelines. Aimed at people already past "which model should I use" and into "how do I make this reliable." Full details here. submitted by /u/camerongreen95 [link] [comments]

00camerongreen951w ago
Databricks CommunityGenie Hub

How to get people to actually choose Genie: train the mindset, not just the buttons

001w ago
Databricks CommunityGenie Hub

DataSmart

001w ago
Reddit

Claude Academy's Using Databricks for Data Analysis - This official documentation from Anthropic on Claude Academy for how to work with Databricks is really underwhelming, so I wanted to share a guide that properly explains everything.

Hey Databricks community! I have been working with Databricks as an admin for over 4 years and I was attempting to follow this official tutorial from Anthropic for how to connect Claude with Databricks https://academy.claude.com/tutorials/using-databricks-for-data-analysis And I encountered multiple significant frustrations with doing it, as the tutorial is very underexplained and outdated for how to get Claude working with Genie One and other functions within Databricks. Therefore, I thought this video on Youtube would be valuable content to share to help solve this problem, for other people that are encountering similar difficulties when getting these two tools connected and building out systems that work for Claude to properly interact with Databricks. Let me know if this video is helpful! submitted by /u/k_kool_ruler [link] [comments]

00k_kool_ruler1w ago
Databricks CommunityData Engineering

Connecting to Snowflake via Workload Identity Federation

001w ago
Databricks CommunityAnnouncements

Announcement | EBOOK: Operationalizing Agent Fleets at Scale

001w ago
Databricks CommunityCertifications

Your exam is being suspended for non-compliance with requirement set by test sponsor

001w ago
Reddit

SDP pipeline parameters

Recently attempted to use pipeline parameters with my SDPs but it seems they are only supported in SQL at the moment. This is a PITA for me using python as I am using the pipeline config as an alternative but this doesn't support updates at runtime. Could you confirm if pipeline parameters are coming to python SDPs and if so, when? submitted by /u/dvartanian [link] [comments]

00dvartanian1w ago
Databricks CommunityData Engineering

Change data feed from a materialized view

001w ago
Reddit

Your Databricks query is slow — what would you fix first?

Imagine you have a Delta table with several billion rows . A query that used to take 30 seconds now takes 8–10 minutes. The query itself looks reasonable, but the table has: A large number of small files Data that is not well organized for the most common filters Several joins Frequent incremental writes You could approach the problem in several ways: OPTIMIZE the table Change the partitioning strategy Use liquid clustering Improve the query itself Change the join strategy Reduce the amount of data being scanned Reconsider how the data is written What would you investigate first, and why? Would you start with the Spark UI/query plan, the table's physical layout, or the SQL itself? I'm especially interested in different approaches rather than one specific solution. submitted by /u/No_Ambition8323 [link] [comments]

00No_Ambition83231w ago
Databricks CommunityGenerative AI

Knowledge Assistant fails to load after indexing; generated AI Search index cannot be reused

001w ago
Databricks CommunityAnnouncements

Community BrickTalk | Databricks Certification: Did You Know?, What's New, and Where We're Going

001w ago
Databricks CommunityGenerative AI

What Are the Key Considerations for Building AI-Powered Applications?

001w ago
Databricks CommunityData Engineeringanswered

Liquid clustering vs. partitioning for a 5 TB Silver table with frequent MERGEs

001w ago
Databricks CommunityData Governance

What information about external agents can be retrieved from Unity Catalog?

001w ago
Databricks CommunityMVP Articles

Unlock the Power: Databricks Genie One vs. Code and Agent

001w ago

Week of Sep 21

62 questions
Reddit

Genie App Builder

What apps are you guys building? submitted by /u/Sea-Glass7015 [link] [comments]

00Sea-Glass70151w ago
Reddit

trigger on_bundle_deploy

In a job bundle, we can specify the trigger on_bundle_deploy in job_runs. Once the bundle is deployed, the job will run automatically. This is really useful for post-deployment tasks like DDLs or one-time table population, such as a calendar. More news and the best articles https://medium.com/databrickscommunity submitted by /u/hubert-dudek [link] [comments]

00hubert-dudek1w ago
Reddit

Preparing for Databricks Champion — looking for guidance

I’m planning to prepare for the Databricks Champion program and wanted to connect with people who have already gone through the process. submitted by /u/Ok-Golf2549 [link] [comments]

00Ok-Golf25491w ago
Databricks CommunityCommunity Articles

Architecting a Medallion Lakehouse

001w ago
Databricks CommunityGet Started Discussions

Balancing Model Agility and Centralized Governance with Mosaic AI Gateway

001w ago
Reddit

UC Berkeley renames football stadium to Databricks Field in sponsorship deal

FTA: Cal Athletics has announced a multiyear partnership with Databricks, a data and AI platform for enterprises, that includes renaming Cal football’s home field to “Databricks Field at California Memorial Stadium.” The San Francisco-based company was founded at UC Berkeley and now will have its name attached to a wide variety of aspects of Cal football, most notably the field at the stadium and the helmets of the players. Sounds like Databricks didn't have to pay any cash upfront for the rights, but is instead giving UC Berkeley stock in exchange. While Cal Athletics and Learfield did not disclose the financial terms or length of the partnership — besides stating that the multiyear partnership will be funded through UC Berkeley becoming a shareholder in Databricks — the deal is reportedly worth $22.8 million for up to 10 years, according to Ben Portnoy of FOX Sports. submitted by /u/gamescan [link] [comments]

00gamescan1w ago
Databricks CommunityDatabricks Free Edition Help

Jobs suddenly stopped running with no errors

001w ago
Reddit

Data Lakehouse Architecture Guide

submitted by /u/codingdecently [link] [comments]

00codingdecently1w ago
Reddit

Use Genie to crate a dashboard to show cost per catalog or workflow

Hi there ! Is is feasible to ask Genie and guide me to create a way of showing databricks costs (we used to do it in azure by adding labels). We want to have it directly in databricks. We use catalogs for each client. Each client has a workflow. Please let me know your thoughts or if you have better suggestions. submitted by /u/francebased [link] [comments]

00francebased1w ago
Databricks CommunityGet Started Discussions

Zempira: dagelijkse capsules voor darmbalans en gewichtsbeheer

001w ago
Databricks CommunityData Engineering

AUTO CDC SCD Type 2 with late-arriving deletes: edge cases the docs don't cover

001w ago
Reddit

Stop Moving Data Into Spreadsheets. Databricks Is Moving the Spreadsheet to the Data.

submitted by /u/TowardsDataExp [link] [comments]

00TowardsDataExp1w ago
Databricks CommunityDatabricks Free Edition Help

Accidentally removed the only workspace admin – Free Edition

001w ago
Databricks CommunityData Engineering

Databricks now deletes unused service principal secrets after 90 days

001w ago
Databricks CommunityData Engineering

(removed)

001w ago
Databricks CommunityAdministration & Architecture

Accidentally removed the only workspace admin – Free Edition

001w ago
Reddit

Customerlake - first impressions, pointers, highlights?

I'm about to dive in, and wanting to set my expectations. Advice from real people appreciated. submitted by /u/KeyMammoth1348 [link] [comments]

00KeyMammoth13481w ago
Reddit

Agents in Databricks

Hello! For some time I have been thinking in AI Engineering and I was wondering what kind of agents I could deploy in Databricks with the entire AI module. Have you implemented agents in PRD? What tasks or value do these agents add to their processes? A while ago I had the idea of creating a PII classification agent, until Data Classification has been released. I would like to hear your experiences and get inspired. Thanks! submitted by /u/Weekly_Marionberry_3 [link] [comments]

00Weekly_Marionberry_31w ago
Reddit

questions on sdp-meta

Planning on adopting sdp-meta , i have few questions. i may test drive it in couple of days but want to check with users who may have adopted it. If we are consuming 100's of tables daily dump via AUTOLOAD , Do i infer schema or should i derive schema for individual table Can i apply any custom logic transformations from bronze to silver other than simple select expr and filters like group by and joins. submitted by /u/Competitive-Fee-4006 [link] [comments]

00Competitive-Fee-40061w ago
Databricks CommunityData Governance

Open-source app on Databricks system tables - looking for people to build it with me

001w ago
Databricks CommunityCommunity Articles

Declaring a Primary Key RELY turned my dbt unique Test Green on Duplicate Data

001w ago
Reddit

Operational Responsibilities in DBX

Has anyone come across job responsibilities that aren't software developers (engineers, analysts, or scientists)? In other words, I'm wondering if there is such a thing that corresponds to a classic DBA position, but for this Databricks SaaS? Now that UC is here with managed catalogs, the technology is looking more like a conventional database platform. In fact the new lakebase is actually a conventional Oltp based on postgres (or LTAP if you prefer the latest marketing). It seems like the ecosystem is complex enough now to support some operational job positions, especially if the data estate is large. Anyone have a position that is more than 50 pct operational?) What about 100 pct operational, like a devops or DBA that focuses on databricks submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36971w ago
Databricks CommunityDatabricks Academy Learners

CIF architecture diagram inconsistent with the description

001w ago
Reddit

Is Databricks Replacing Power BI and Tableau in the Enterprise?

Why did Databricks get into the BI tooling space & what does it mean for Databricks customers using Power BI, Tableau, and other established BI tools? In this 3-min video, David Meyer (SVP @ Databicks) shares how Databricks can help both provide advance decision making capabilities and lower costs, as well as where some of the older tools fit in your stack. This was recorded in front of a live audience at the Databricks Data + AI Summit 2026. Hope you enjoy it, and would personally love to hear your stories, both where AI/BI Dashboards from Databricks has been able to take on more of your BI workloads, as well where you still find gaps that need to be addressed! submitted by /u/JosueBogran [link] [comments]

00JosueBogran1w ago
Databricks CommunityCommunity Articles

Learn Databricks AI Agents | Build, Deploy & Evaluate on Databricks Platform

001w ago
Databricks CommunityData Engineering

Best Architecture Approach for Databricks Multi-Cloud and Multi-Region Deployment

001w ago
Databricks CommunityGet Started Discussions

Configuring Genie: practical steps for more reliable answers

001w ago
Databricks CommunityData Engineering

Databricks Workflow Task Failure - Custom Error Messages

001w ago
Databricks CommunityCommunity Articles

Tech summit FY 2027

001w ago
Reddit

Databricks Usage in real world?

Hi all, I want to understand how AZure Databricks is actually being used in real world data engineering teams. For those of you working with Databricks, how do you use it in your current job? Likee, do you write most of your code in notebooks and publish those notebooks through Git? Or are you using Asset Bundles for ci/cd and deploying jobs and workflows through different environments? I’m trying to get a clear picture of how ADB is used in actual companies. Help this buddy out. TIA!!! submitted by /u/Wise_Beat_7035 [link] [comments]

00Wise_Beat_70351w ago
Databricks CommunityGenerative AI

DeepSeek V4.1 Flash returns workspace ITPM rate-limit error on fresh/idle workspace

001w ago
Databricks CommunityData Engineering

Genie One Foundations: What Data Teams Should Get Right Before Rolling It Out

001w ago
Databricks CommunityBrickTalks TV

Bricktalk Recording | Real-Time Data & AI: Tripwise Demo

001w ago
Reddit

Azure devops yml deployment of asset bundle

I need help on deploying asset bundles from Azure Devops. Where does the azure-pipelines.yml live, what parameters are required in the azure-pipelines.yml, does anyone have a basic example of it working as I'm completely lost in the complexity of it. Thanks. submitted by /u/GardenShedster [link] [comments]

00GardenShedster1w ago
Databricks CommunityData Engineering

org.apache.iceberg.connect.IcebergSinkConnector to sink data from kafka to databricks - large volume

001w ago
Databricks CommunityTraining offerings

Demo Material: Building and Orchestrating an Agentic App - AI Agent Fundamentals Plan

001w ago
Databricks CommunityDatabricks Academy Learners

Where to find Notebooks for course 'Advanced Techniques with Apache Spark Declarative Pipelines'

001w ago
Databricks CommunityData Engineering

Anyone migrated from legacy CDF to Auto CDF yet?

001w ago
Reddit

Lakebase handles failover within a region, but cross-region disaster recovery is still in Private Preview. How are people covering that?

I was reading through Databricks' breakdown of what Lakebase takes over on the ops side. Patching, scaling, in-region failover, and point-in-time restore with 2 to 30 days of history all run automatically, while cross-region disaster recovery is Private Preview on AWS only, with manual failover and recovery procedures the customer owns. For anyone running Lakebase behind an app or for agent state, how are you handling a region outage today? I can see people copying scheduled snapshots elsewhere, replicating out to a Postgres instance outside Databricks, or accepting the risk until DR goes GA, and I'm curious which one teams have landed on. I'd also like to hear whether the built-in PgBouncer has held up under spikes in connections from Databricks Apps, since that's the other thing I'd want to test before putting a production workload on it. submitted by /u/InsideDebt6345 [link] [comments]

00InsideDebt63451w ago
Databricks CommunityAdministration & Architecture

Managed Disaster Recovery

001w ago
Reddit

Do you actually partition your Delta tables, or has everyone moved away from it?

keep seeing conflicting advice around partitioning Delta tables, and I'm curious what people are actually doing in production. I understand the general idea of partitioning, but in practice I'm finding it harder to decide when it's genuinely useful versus when I'm just creating more small files / directories for the sake of it. For example, imagine a table with hundreds of millions of records that's mostly queried by date, but also has filters on things like customer, region, status, etc. Would you partition by date? Leave it unpartitioned and rely on Delta's optimizations? Something else? And for people who have actually dealt with this in production — what made you change your original approach? I'm particularly interested in “we did X, it caused Y problem, so we changed it to Z” stories. Those are usually more useful than the textbook explanation. submitted by /u/Delulu62134 [link] [comments]

00Delulu621341w ago
Databricks CommunityCommunity Articles

USE CONNECTION bypasses your MCP Service's Tool Selection and Policies

001w ago
Reddit

UC ingestion from a SQL server

I'm creating a UC catalog. Lets say I want to ingest a SQL database that contains 50 tables, parented by five meaningful SQL Server schemas. (Fact.WhateverThing, Dim.WhateverThing, Logging.Whatever, CONFIG.Whatever, and VIEWS.Whatever). I plan to name the catalog in a conventional way ("sales_dev",) but I'm stumped for naming UC schemas and UC tables. Since the Dim and Fact tables need to be joined frequently, I'm assuming all these tables should be in a SINGLE unity catalog schema, yes? Should I simply move all five of these SQL schemas into a single UC schema (maybe into a single UC schema named "bronze_from_sql")? At that point how would I name the tables? I'm assuming I would be forced to do something like so: "fact_whatever_thing", "dim_whatever_thing", "logging_whatever", "config_whatever", "views_whatever"? It strikes me that UC doesn't have the same concept of schema that we have in SQL server. A schema in UC is actually referring to a whole database. Whereas in SQL it can be used to organize tables within a database. Edit: decided to publish sql server tables to volumes (in a bronze layer). Will just use delta format, and free-form folder-and-table names. Is very liberating not to conform to the constraints of the structured table names in UC submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36971w ago
Databricks CommunityCertifications

My Databricks Certification exam suspended randomly

001w ago
Reddit

Databricks Acquires Row Zero, Bringing Live, Governed Spreadsheets to Genie

submitted by /u/sai-nageshwaran [link] [comments]

00sai-nageshwaran1w ago
Reddit

Problems with UC MCP connection

We built a supervisor agent where one of the tools is the native databricks MCP connection (which is in beta). Today, the agent doesn’t work. Removing the MCP tool fixes it, so it’s definitely that. The agent response just says “received an invalid and unexpected value from the API: undefined” Anyone else having problems suddenly with the MCP server? I’m able to hit the target MCP URL locally no problem. And even test the UC MCP in databricks successfully. So it’s something about the handshake between MCP and the supervisor agent submitted by /u/pboswell [link] [comments]

00pboswell1w ago
Databricks CommunityData Engineering

Operationalizing Lakeflow Connect:Handling Upstream Schema Evolution & Historical Backfill

001w ago
Reddit

Laya off the benchmark: can a zero-shot decision model route real SQL traffic?

Weekend-ish experiment on Databricks. One table, TPC-H orders, 15M rows, living in two places at once: Lakebase (Databricks' managed Postgres), a continuously synced copy with a btree index on the key. Sub-100ms point lookups, useless for a GROUP BY over 15M rows. Delta behind a serverless SQL Warehouse. Great at scans and aggregations, slow at fetching one row. The synced table is the nice part: native continuous Delta to Lakebase Postgres (needs a PK + CDF), so it's one dataset under one Unity Catalog, replication handled by the platform instead of a homemade pipeline. Then I put laya (convaiinnovations/laya, a non-autoregressive zero-shot decision model, vanilla, no fine-tuning) behind a FastAPI Databricks App. It reads each query and picks the engine. MLflow traces input to decision to execution. I submitted by /u/Limp-Park7849 [link] [comments]

00Limp-Park78491w ago
Databricks CommunityAdministration & Architecture

Serverless DNS failures after trial upgrade (AWS us-east-2)

001w ago
Databricks CommunityGet Started Discussions

From Databricks POCs to Building a Data + AI Community

001w ago
Databricks CommunityLakebase Discussions

Lakebase Data API (GCP) returns jwk not found for valid service principal tokens

001w ago
Reddit

Databricks Acquires Row Zero, Bringing Live, Governed Spreadsheets to Genie

submitted by /u/orangepunc [link] [comments]

00orangepunc1w ago
Databricks CommunityAnnouncements

Announcement | Genie One MCP: Give any AI Agent the Right Business Context

001w ago
Databricks CommunityCommunity Articles

Tags no Databricks: Governança, FinOps e MLOps

001w ago
Databricks CommunityData Governanceanswered

How to override ABAC policies

001w ago
HackerNews

Databricks Acquires Row Zero, Bringing Live, Governed Spreadsheets to Genie

30ms5121w ago
Reddit

Genie App Builder, App Spaces, and Micro Apps released to beta

Yesterday (Sept 23), the app features all of us eager for since they were previewed at Summit were released. In testing, it was not available in all workspaces (cloud and region available differ). Looks like you have to use App Spaces to use these features. App Spaces is a collection of app(s) with some associated boundaries. I'm thinking like if you had a group of apps for a specific team with same users and data and developers, but since brand new don't know if that is exactly how they work. App Builder - A simple UI front end for non-technical users to build an app using Genie. Will be curious to see how this is different than the chat Genie Code interface. Micro Apps - This is the feature I think most of us wanted, the ability to scale to zero, and only pay for actual app consumption with no minimum spend Really excited to see these in the wild. submitted by /u/ecp5 [link] [comments]

00ecp51w ago
Databricks CommunityData Engineering

Genie Use Cases & easy adaptibility

001w ago
Reddit

DQ Management

Hi, we have a lot of data products, how do u design your data quality, and what tools do u use? Idea is that we use dqx, we define rules (ERROR, WARN), and enforce that in pre write of the data, which is to be honest a new thing. Then i would also like to have anomaly checks, where i could compare stats of current batch run to the previous one. How do u design your dq around dqx? And what else do u use? What "flow" do u use? submitted by /u/ptab0211 [link] [comments]

00ptab02111w ago
Databricks CommunityAdministration & Architectureanswered

Access to Lakehouse Real-Time (RT) in the preview tab

001w ago
Databricks CommunityGet Started Discussions

🚀 Databricks AppQuest Quest 5 – Intelligent Knowledge Base with RAG

001w ago
Reddit

Conditionally retrying a task on Lakeflow Jobs

I've entered a new company which uses DAB for all orchestration, and my first task has been to automate certain pipelines that fail due to data unavailability so they retry the fetch later instead of requiring manual runs. Issue is that data fetching errors are one of the many kinds that can be raised during the tasks, and it should be the only retryable one. I'm new to DAB and I've been diging into docs and forums, but I cannot find a native approach or non-patchy workaround to only retry on certain conditions (like based on exit codes). So far, I've come with these candidate solutions: - Just retrying everything (which doesn't makes sense cause if any of the other errors occur, all attempts would fail) - Failing only for the data fetch and propagating fatal errors to the following task, which would check state and fail if errors occurred - like via notebook return values (I don't like this because couples orchestration into business logic, even in a distinct layer of the one that generates the error, plus, the task that would appear as failing would be the successor to the fetching task, which would be shown as successful - not to speak about the hundreds of tasks that would need source code changes) - Failing only for the data fetch and adding an intermediate if/else task that checks if (via notebook return) the previous task propagated and error and stopping the job in such case (this is for now the most likely, as I can keep everything at orchestration level, and it is a bit more explicit even if the fetch task appears as successful, though it still feels like a workaround and tons of jobs would have to implement this new intermediate task) A couple of clarifications: - I cannot implement the retry within the task as fetches would be repeated after some time, and we don't want to have resources running and being billed just for waiting - Everything is done via Lakeflow jobs and with Databricks computer options (even if Spark is not needed) - Most of the fetching occurs by calling external FTP servers, data doesn't directly arrives to our infrastructure like to trigger the pipelines via events Thank you all Edit about explicit behavior: - On data fetch related error, retries occur - On other errors that can come from the task, the whole job fails without retrying anything - On success, next task executes - If failing due to a non retryable error, or due to maximum retry, the job fails without executing downstream tasks submitted by /u/Xandor19 [link] [comments]

00Xandor191w ago