Ce que la communauté demande.
Discussions récentes de r/databricks et du tag [databricks] de Stack Overflow : problèmes concrets, questions d'intégration et cas particuliers à connaître.
This week
57 questionsRecording | BrickTalk: Mastering Databricks Genie Capabilities
Endpoint Error
Databricks Certification Exam Suspended - Databricks Certified Data Engineer Associate
Hospino : Hospital Management Software
Confused new user: where to find the notebook for free course?
Questions about "Store OpenTelemetry traces in Unity Catalog"
Using autoloader with multiple object types in load path
Databricks Dashboard - Scheduling Questions
( FYI, I did write the initial draft of this message by myself, and then later asked AI to proof read it and later pasted the same here, Apologies in advance!) I recently developed and published a Databricks dashboard for a personal project, and I want to learn how the scheduling feature works. For context, I have read the Databricks documentation, Manage scheduled dashboard updates and subscriptions | Databricks on AWS , but I couldn't find answers to the questions below. Databricks Dashboard — Question Set 1 I published the Databricks dashboard using Individual Data Permissions . I then added another person as a participant, created a schedule, and subscribed to the schedule from both accounts. Let's call them Person A (publisher + subscriber) and Person B (subscriber). Questions a) Assume Person B works in the Sales department and should only be able to see Sales data. Since the schedule was created by Person A, when the scheduled run occurs and Person B receives the email notification they subscribed to, will the PDF attachment contain all departments' data , or will it contain only the data Person B is permitted to see? b) After the scheduled run, if Person B opens the dashboard using the link in the email notification, or accesses the dashboard directly through Databricks, what data will they see? Will they see all department data , or only the data they are permitted to see based on their individual data permissions? My understanding is that when Person B visits the dashboard, it does not automatically refresh just because a scheduled dashboard update has occurred. I may be misunderstanding how this works, so I'd appreciate some clarification. Databricks SQL Alert — Question Set 2 I also created a SQL alert to notify me when a new entry is added to a dataset within the last 24 hours. Please note: The screenshot below is not from my actual example. I am only including it so that new developers like myself can relate to what I am referring to. enter image descr […truncated]
transformWithState might be causing 'Module not found'
Streaming Doesn't Mean Your Compute Needs to Run Forever
Evaluating AI Agents Live at the Grounded Reasoning Cup
Announcement | How Databricks Genie Code Automated 90% of Data Ingestion for a Major Railroad
DQX Forge - Extension for DQ in Databricks
Genie Ontology Benchmark
request for Databricks Certified Associate Developer for Apache Spark 3.0 free voucher
☕ Qué es un CLUSTER en Databricks | Tipos, DBR, Serverless, Costes, DBUs y Spot Instances 🇪🇸 🚀
Databricks App Access Issue - Free Tier Platform
Estimating Databricks Annual Implementation and Licensing Costs in 2026
Vector Store update stale in Syncing status even the actual sync task is done.
Bedrock and Genie
Learn Databricks Agent Bricks | Build Enterprise RAG Agents
CUSTOMER STORY | From high costs to near real-time insights with Zerobus Ingest
Microsoft Fabric One and Databricks Unity Catalog — Bidirectional Access Without Data Movement
Possible Incorrect Answer Marking in Databricks Academy Assessment - Agent Evaluation on Databricks
Databricks delta tables and Iceberg
DATABRICKS.SQL and DATABRICKS.TABLE Functions Debuts in Microsoft Excel
Solution Accelerator Series | Overall Equipment Effectiveness
🌟 Community Pulse: Your Weekly Roundup! August 10 – 16, 2026
how to enable column level lineage in databricks
Auto SCD API Tombstone Garbage Collection
Not Able to Remove Courses
Announcement | Taking AUTO CDC to the Next Level in Databricks
Open Sharing protocol of Databricks Apps and SaaS sharing
I want to understand the open sharing protocol, how it is sharing the apps, what artefact is shared when app is shared, is it sharing the app files from Workspace path directly? Secondly, how the SaaS is shared and what artefact is shared and how? Could not find documentation. Thanks
Identity Columns Best Practices for Databricks Lakehouse
Cannot work with workspace
DABs Migration Guide: Terraform to the Direct Deployment Engine
Not Able To Log In - Free Edition
Azure Databricks- Unable to enable datatbricks admin role to user to have access to system.serving catalog
This looks like a strange situation that i am facing here i have created a databricks with existing user in my personal account the user already have global administrator privilege however after i signed into datbricks, this user(personal id) appeared as worksapce admin rather a account admin. i realized this issue while attempting to access 'system.serving' catalog/table where the databricks not allowing logined user to access/grant permission to the system catalog. As per the Microsoft documentation, when we create a databricks workspace / login to workspace, it initially provide the privilege as Account admin, but it seems account admin privilage is not granted. Is there any solution to fix this? #Expected: my currnent user should have databricks admin rather workspace admin access
Genie Agent Volume Attachment (Beta) Docs mentions multiple formats but images aren't supported
Account Console environment for demonstrating workspace management capabilities to clients ?
Child-Parent Partnership
S3 LIST costs on high commit rate Delta tables: is there a start-after option on Databricks runtime?
ML Training low File I/O and Throughout
Databricks Community Contest | Genie-Powered App Challenge
where to find the lab to practice
Azure Databricks Serverless Compute Unable to Connect to Azure SQL MI Using Failover Group FQDN via
🇪🇸 ☕ ¿Cómo organiza Databricks tus datos? | De Unity Catalog a Delta Lake ⚡
CUSTOMER STORY | Albert Einstein Hospital cuts clinical query time with Genie
Beyond Prompting: Production-Ready Image Generation and Visual Asset Management on Databricks
How to Block Databricks Genie Usage When a Budget Limit Is Reached
Configure Databricks Access to Keyvault with Azure Role Based Access Control (RBAC)
Generative AI Engineering Pathway/Agent Evaluation on Databricks/Quiz
I want to help me with registration key
Databricks Genie Cost Control: How to Set Budgets and Block Usage
If statement in DAB YAML file support
Unity Catalog Lineage APIs – Are these REST APIs publicly supported?
Is the Semantic Layer the Real Challenge With Databricks Genie?
Last week
93 questionsOracle CDC Ingestion Pipeline - Schema Exploration finds 0 tables when DB_DOMAIN
Databricks as a Semantic Engine: Why the Semantic Layer Was Never Enough
test azure upload file
Exploring Databricks Instructed-Retriever-1 Through a Data Engineering Use Case
Exam issues due to Webassessor maintenance
Issues with Custom Agents on Free Edition
Are We Entering the Context Engineering Era?
Request to Merge Two Databricks Community Accounts
Visual, step-by-step curriculum for mastering DABs and Infrastructure as Code
SQL Query billing
Billing
API Get Metadata Registerd Models
How to use AI for photo filter during registration like OwnMates?
Omnigent Meta-Harness
facing issue in llm models
Excel Add-in Regression - Sign-in opens external browser window instead of embedding in task pane, r
Databricks Raises $5B at a $190B Valuation
why micro-batching matters so much in Databricks Auto Loader and Structured Streaming
End-to-End Streaming NLP Pipeline with GDELT, Azure Data Factory, ADLS Gen2 and Databricks
How to extract data from SAP to Databricks?
End-to-End Streaming NLP Pipeline with GDELT, Azure Data Factory, ADLS Gen2 and Databricks
Will Databricks Ever IPO?
--- top comments --- [jethronethro] Does Databricks need to IPO? If so, why? [eyehurtsme] They have 0 reason to IPO at the moment
Lakebase synced table doesn’t recognize Auto CDF on a SDP materialized view
Oh no, not another one! Databricks buys Electric
Gear Up: The Next Databricks Community Challenge is Almost Here!
Can I use a corporate voucher on my personal Databricks account? (Plus: discount type & expiry)
Solution Accelerator Series | R&D Optimization With Knowledge Graphs
Data quality Lineage Root cause analysis
How to call a SQL Server stored procedure using pymssql from Databricks Serverless Compute?
I have a Databricks notebook that currently uses pyodbc to connect to SQL Server and execute stored procedures. I need to migrate this notebook to Databricks Serverless Compute , so I am looking for an alternative to pyodbc . I am considering using pymssql instead. What is the correct way to connect to SQL Server and execute a stored procedure using pymssql from a Databricks Serverless Compute environment? As of now this is the code we're using: def exec_stored_procedure(stored_procedure, json_data): try: conn = pyodbc.connect(connection_string) cursor = conn.cursor() cursor.execute(f"OPEN SYMMETRIC KEY {Symmetric_name} DECRYPTION BY PASSWORD = '{Symmetric_key}'") cursor.execute("{CALL " + stored_procedure + "}", json_data) conn.commit() except pyodbc.Error as e: print("PyODBC error:", e) except Exception as e: print('An error occured: ', e) finally: try: cursor.execute(f"CLOSE SYMMETRIC KEY {Symmetric_name}") except pyodbc.Error as e: pass except Exception as e: print("An error occurred while closing symmetric key:", e) try: cursor.close() except pyodbc.Error as e: pass except Exception as e: print("An error occurred while closing cursor:", e) try: conn.close() except pyodbc.Error as e: print("PyODBC error while closing connection:", e) except Exception as e: print("An error occurred while closing connection:", e) What would the equivalent implementation using pymssql look like, and are there any additional requirements or limitations when using pymssql with Databricks Serverless Compute? My goal is to replace pyodbc while keeping the existing SQL Server stored procedure logic unchanged.
NetSuite JDBC Driver 8.10.190.0 - Databricks support
Difference in exam content between the PT-BR and English versions - Databricks Certified Data Engine
Object metadata
X_NHC_CONTROL_PLANE_UNREACHABLE
🌟 Community Pulse: Your Weekly Roundup! August 03 – 09, 2026
Skipping malformed records when reading Avro-files
Learn Databricks Lakebase: managed Postgres for apps, agents & real-time data
From 90 Minutes to 3: We Turned Our Genie Space Into an Employee
Open-sourcing Metals v2: Databricks' Java and Scala language server
Databricks streamlit app with write back capability and audit trail display
Learn Databricks Genie: 5 courses to go from curious to certified
LLMOps for Data Scientists and AI Builders: A Quickstart on Databricks
Databricks Community Fellows – July 2026 Recap
Unity Catalog Metric Views to be accessible to Custom Apps outside of DBX environment
Data Drift Metrics
Integrate genie workspaces automatically based on connection.
Migrate your Dashboards to AI/BI with Genie Code
Electric is joining team Neon at Databricks
Change Data Feed in Databricks Delta – How to Process It the Most Efficient Way
Databricks App architecture with AppKit - agentic app with governed write-back functionality
Electric Is Joining Databricks
Error DELTA_CATALOG_MANAGED_TABLE_UPGRADE_WITH_OTHER_PROPERTIES during catalog commit upgrade
Declarative Until It Isn't: Four Sharp Edges of Lakeflow Declarative Pipelines
How to remove/delete a Relationship Graph (Semantic Model) from an AI/BI Dashboard?
Electric is joining team Neon at Databricks
Neon acquires ElectricSQL to build better data syncing for agents
Track Secrets Access
Request for Name Correction on Databricks Certification
Why is the default auto-termination for serverless interactive notebook compute 60 minutes?
Electric(SQL) Joins Databricks
Multiplex Streaming, Delta Sinks, and Iceberg Reads with Databricks
Databricks Watermark-Based Incremental Ingestion
I benchmarked the new Databricks Lakehouse RT for billion-record tables
Electric is joining team Neon at Databricks
--- top comments --- [randombetch] Sick. Neon rocks. [randombetch] Sick
Electric Is Joining Databricks
Synced Tables - Partitioned Tables
Announcement | Databricks Completes Acquisition of Panther: Accelerating the Security Lakehouse Era
Databricks EMEA Learning Festival | September 29-30
Unable to Enable Unity Catalog – Azure Managed Identity Credential Not Found
Unity Catalog service credential get_token rejects api:// scope format — "not a valid URI"
AI/BI Dashboard pivot export: Excel
Multi-fact Star Schema patterns in Databricks
Unexpected behavior of Delta VACUUM – need explanation
Upcoming Community BrickTalk: Mastering Databricks Genie Capabilities
Disabling Change Tracking and enabling Change Data Capture in SQL Server Lakeflow
Is VARIANT supported in Databricks-to-Open sharing?
How to manage SQL queries for business data extraction in Databricks?
Azure Databricks Default Package Repository with Azure Key Vault-backed Secret Scope
From 40 Minutes to 8 minutes: Why We Dropped MERGE in Our SAP BW to Databricks Gold Layer
Solution: Simplify Genie Agent Instruction Updates Across Multiple Spaces
Creating Databricks agent
Triggered vs. Continuous Mode: A Deep Dive into Serverless Lakeflow Spark Declarative Pipelines
Google Drive ingestion pipeline failing – “Google Drive file system is not enabled”
How to Build, Test, and Ship SQL Pipelines on Databricks
Wrong Exam Selection – Request to Change from Data Analyst Associate to Data Engineer Associate
Lakehouse Monitoring Solution
How should schema evolution be handled across silver and gold layers in a medallion architecture?
Building a Production LangGraph Agent on Databricks - NorthStar Brand Copilot
Databricks Community Champion - July 2026 - Emma Stowell
Databricks Cost Optimizer: Audit Spend with Codex or Claude Code
Databricks Data Mesh Best Practices: A Practical Implementation Guide
Databricks Data Mesh Best Practices: Practical Implementation Guide
Learning Series | SQL Programming and Procedural Logic
Databricks Support #00984257 - Password Reset Email is not working for me
Week of Aug 3
50 questionsRT Lakehouse - impossible?
Change Data Feed on Materialized Views Why I Think This Is More Than an Incremental Processing
Request for Databricks Associate Engineering Exam Voucher
From Spreadsheets to Insights: How Genie One Transforms Excel
Getting on-premises SQL Server data into Databricks: the networking is the hard part
Feedback on Deploy Workloads with Lakeflow Jobs SPL
Is it safe to expose JWT in Databricks Job?
Databricks Knowledge Assistant
From Agent-ready to Audit-ready: The Evidence Layer of AI Governance
Monorepo vs Multi-repo for Databricks Asset Bundles: A Decision Framework
Building Custom Apps on Lakehouse
Managing AI Coding Costs at Scale
--- top comments --- [extr] I would be really curious to hear from devs at Databricks what the experience of development is like internally. I work at a small startup with essentially unlimited AI spend budget - the entire point is that I should be turning to it at every opportunity since our human labor is so expensive relative to tokens. So generally it's like: - Spend most time prioritizing/discussing what to do. - Once that's agreed, use Fable 5 High + 5.6 Sol XHigh come up with a design + plan. Agree on the high level plan. (Usually this just comes down to choosing where the change belongs on the spectrum between minimal patch <-> full redesign) - Use Opus 5 or Sol Med to execute - Auto-fix bugs and CI until green + thermonuclear review skill x3. - Manual interrogation of change/nits - Come up with QA plan and have Codex Computer Use execute on it - Manually spot check the final result (usually a sizable diff, thousands of lines, complete feature E2E, etc) I probably spend like $80 a day at least but I produce the output of 3 or 4 2022 engineers and probably at better quality. So it's easily worth it. Would I save money by switching to GLM 5.2 and such...perhaps? IDK. At our scale it's not worth the time spent building the eval harness to actually understand the performance tradeoff. [lbriner] There are a surprising number of articles like this along the lines of, "we started using AI tools and ended up spending millions per year". On what planet do people start paying for things without keeping an eye on the costs and no-one notices until you have spent a crazy amount? I don't understand. You are either paying a fixed amount which you are happy about in-advance or you are PAYG in which case you would ballpark how much it costs. Otherwise it reads a bit like a fake problem, because it didn't really happen, you just foresaw it (as you should) and added a few guide rails. [sashank_1509] I suspect that when it comes to hard complex software products, you’re better off ignoring agents and doing “trad coding”. What you lose in short term speed you gain in manageable complex codebases. If you have a 500k line codebase and even > 50% is written by agents, you are in a world of pain that won’t justify the costs longer term. Now of course, there are products that just involve lots of code but are not actually complex. This is generally the project with like hundreds or thousands of features but most of the features are separate and don’t actually interact in complex ways. Think a task management app with hundreds of features like calendar, email integration etc. there I think agents gives you more bang for the buck. Just my thought, using agents at work. [platinumrad] Careful. If you admit to using models that weren't trained by OpenAI or Anthropic then you might hauled in front of Congress: https://www.scmp.com/news/china/diplomacy/article/3362616/us... [dgellow] What I take from this is that models are already commoditized, and it’s pretty clear nobody has a moat: routing for the models, they can be swapped whenever new models are released, AI labs will have to continue to run on the treadmill non stop or be replaced. Long term I cannot imagine that business will be high margin. Routing for the harness, so anything that differentiate a provider vs another isn’t exposed to the user and isn’t too relevant. One more datapoint for the thesis that OpenAI and anthropic aren’t viable, sustainable businesses, and cannot justify their $1T valuation and the level of compute commitment (reminder that OpenAI committed to >$750B in infra spending for 2030)
Synced table pipeline fails with permission denied for database
Unity Catalog in Microsoft Azure Government
Plug & Play: Zerobus Ingest Now Supports Apache Kafka® Compatible APIs (Beta)
today is Aug 7th, are there databricks certificate exam vouchers available?
CUSTOMER STORY | How Bupa Australia Unified Health Data and Accelerated Analytics 8x with Databricks
Looking for recommendations
available virtual machines in databricks
How Can I Analyze Gaming Website Traffic and Search Trends with Databricks?
Support Ticket Unanswered – Reschedule Issue Due to Delayed Support Response
Building an Agentic HR Front Door on Databricks
DABs: immutable_folder
How to create an image from a cluster so that the compute environment can be replicated
Requesting free/discounted voucher for Data Engineer Associate certification exam
Announcement | Introducing OfficeQA Pro V2: Grounded-Reasoning Benchmark for Enterprise Data
Haven't received a project expert badge
Mono Repo or not Mono Repo - that is the question?
Ticket Escalation: Request for Certificate & Webassessor Name Correction (Case #00980951)
Do you see a future where enterprise MDM platforms like Informatica MDM could be fully replaced ?
Announcement | Take Insights Anywhere with Genie One on Mobile
Dashboard Security Control
completed modules but didn't receive coupon
CREATE is not allowed: Cannot CREATE the Streaming Table Serverless Generic Compute
NetSuite connector (Lakeflow Connect) — transactionline refresh takes 3.5–4.8 hrs regardless
Video Recordings of Previous Webinars
Databricks Partner
Announcement | Unity AI Gateway is Generally Available
Load Testing Databricks SQL Warehouses with JMeter — Part 3: Running and Analyzing Results
Load Testing Databricks SQL Warehouses with JMeter — Part 2: Concurrency and the Test Plan
Load Testing Databricks SQL Warehouses with JMeter — Part 1: Inputs and Configuration
CUSTOMER STORY | Worldpanel by Numerator Cut Reporting from 12 Days to 1 with Lakebase on Databricks
[PARTNER BLOG] Evolving Metric Views: YAML, UI, Genie Code & Materialization
Building a Visitor Data Pipeline for Digital Membership Card
Pi's Minimalism Is Its Advantage
--- top comments --- [paldepind2] Lots of praise for Pi in this thread, so I'll offer up a diverging opinion. Given all the hype, I was a bit underwhelmed by Pi. It definitely has some good ideas around customization, but it annoyed me in many little ways. For a program that's minimal it sure takes a long time to start up, the standard C-p and C-n bindings don't work, it doesn't follow the XDG Base Directory Specification and just pollutes my $HOME directory. I think there's still space for another harness that's 1/ open-source, 2/ written in a fast compiled language (Rust, Go, etc.) and scriptable in a simple (aka non-JS) scripting language (Lua, etc.), 3/ less opinionated and more sensible so things like XDG isn't a WONTFIX. [pavo-etc] I've had a lot of success running Pi on my server in headless mode and wrapping it in an XMPP client. This means I can talk to it wherever I can access XMPP (everywhere). It also mean agents can talk to each other when they need to. They've got a shared wiki they interact with and github issues as their todo list. I am running several named pi instances in parallel in their own user account on NixOS, so they can install whatever they want in ephemeral shells and I never need to worry about their env. The agents can spin up new enabled XMPP agents if I request it, though for now I've only needed a few since I'm not doing too much in parallel. My Pi is very vanilla, only my own XMPP wrapper and pi-subagents extension for anonymous subagents. Using it primarily with Deepseek v4 Flash for chipping away at coding tasks or server maintainence while I'm AFK or in transit. NixOS is the key to all of this, since agents can interact see the whole server config, make changes and run compile-time checks before actually deploying. It also means that even if they do mess up I can always revert. [malisper] Having tried all the coding harnesses, I find that using Pi is exactly like using Emacs. For anything you want to build you can ask your agent and it will build it. There's tons of existing code to help you configure it. At the same time half the code is buggy, UI elements will try to overlap one another, and you'll periodically get crashes. If you're willing to put in the work to master the learning curve and push through the issues, it can be a great tool: https://i.sstatic.net/7Cu9Z.jpg [mft_] I’m going to go against the grain and say that Pi is a little too minimal by default. The emacs comparison is interesting, but default Pi is like emacs that can load files but you’ve got to extend it manually to save or search within a file. I’d argue that there’s a minimal set of functions that a coding harness needs to just enable a model to get stuff done, and they shouldn’t be an extra effort to set up. (Oh-my-pi exists for those of a similar persuasion.) [astrobiased] Pi does one thing that I love, developing a tool that has minimalism where it's easily configurable with good documentation. The leads to new use cases that the the author(s) would have never dreamed of. The organic growth process of the Pi ecosystem has been fascinating to observe. It's one of the reasons why Pi has become one of my favorite coding agents to this day, flexible beyond personal uses and extensible to larger environments. IMO, I view it more than a coding agent, it's a coding agent platform with powerful extensibility.
