What the community is asking.
Recent threads from r/databricks and Stack Overflow's [databricks] tag — practical pain points, integration questions, and edge cases worth knowing about.
This week
32 questionsBuilding an Enterprise RAG Chatbot with Databricks Mosaic AI and Vector Search
New Dashboard for monitoring Genie Usage and Costs
Dashboard for Genie Cost and Usage tracking
How to get started with Genie Ontlogy?
Best Practices and Questions for Configuring High QPS AI Search Endpoints in Databricks
CREATE CONNECTION - Support for Community Connections?
Can {{input}} from a For Each task be used as a dashboard subscription destination_id
Supervisor Agent or Knowledge assistant not available
From Experiment to Prod: LakeFlow Spark Declarative Pipelines Testing Blueprint
Serve Tableau Reports Directly from Databricks SQL
Generative AI Certification
Solution Accelerator Series | Measure Ad Effectiveness With Multi-Touch Attribution
Genie Spaces (Agents) in DABs
Govern AI Spend at Scale: A Data-Driven Approach to AI Governance | Webinar
Critical Genie Agents Issue – Catalog Rename Causes Tables and Joins Loss
Advanced Learning Festival: 15 June - 06 July 2026 - Voucher not received
Free Certification voucher
Support Multiple Tasks DAG Inside a `for_each_task` Iteration
Omnigent: The Control Layer Enterprise AI Needs
What happens to Databricks notebooks and jobs when a user's Microsoft account is deactivated?
Announcement | What happens in the milliseconds after you tap pay
[CUSTOMER BLOG] How S&P Global Energy Made Its Structured Data Estate Conversational with Databricks
When Milliseconds Matter: Databricks Powers the Next Generation of Operational Workloads
Can I connect Fabric Data Agent With Databricks Genie One as external connection
Can I connect Fabric Data Agent With Databricks Genie One as external connection
Exam suspension due to non-complaince reported. Request for review and reschedule
Enterprise Unity Catalog RBAC model in Databricks
unable to open the vcorum labs though I have the subscription
VOID column inside STRUCT fails to cast to VARIANT
Issue with Notebook Scrolling Position While Switching Between Multiple Notebooks
Deactive Admin user
Does user_identity in system.access.audit reflect the actual logged-in user for Service-Principal-mo
Last week
100 questionsDatabricks Free Edition – No Workspace Available
Free Edition account active but workspace missing ("You are not a member of any workspaces")
What Is a Spark Shuffle? A Simple Visual Explanation ⭐
Databrick Genie
Enterprise-level Unity Catalog governance
DLT driver GC pressure during a large initial hydration. Is a bigger driver the only lever?
Data contracts on Databricks: what does a working MVP actually look like?
Genie Ask in CLI
🚀 How does Databricks process terabytes of data in just minutes?
Alteracao de prova ingles para português
Databricks Lineage: Why It Matters More Than Ever
Use of Genie agents for ETL migration
2026 State of AI Agents: Enterprise Insights on Building AI
What happens to with SCD2 tables if the source pipeline performs a Full Refresh?
can i deploy a metric view using DABs
One Pipeline, Any Destination: The ForEachBatch Sink in Spark Declarative Pipelines is now GA
Community Custom Connector - Defining Non-Serverless Job Compute Runtime?
CONNECTION Permission Issue - Can't use existing connection for Community Custom Connector
Auto-Termination Did Not Trigger on Production Cluster Despite 20-Minute Inactivity
Open Source Doesn't Mean You're Free. Managed Service/Closed Source Doesn't Mean You're Locked-in
Databricks AMER Learning Festival | Virtual Training
Lakebase Continuous Sync: Why My Synced Table Stayed Empty Despite a Healthy Pipeline
Databricks Apps (Streamlit) - How to Implement Proper Logout Functionality?
Reducing Time to Help in the Databricks Community
'Schedules' feature in Genie One
UC Permission Automation Strategy
Keep a PySpark Test Flood Out of Your Coding Agent's Context Window
Databricks & Microsoft Expand Partnership to 2030s: Architectural Impact on Azure Databricks
Pipeline still needs USE SCHEMA on an old schema it no longer writes to
Announcing General Availability of the Lakeflow Spark Declarative Pipelines Kafka Sink
Has anyone prototyped Databricks Lakebase or deployed it in a production environment?
AgentOps on Databricks: Operating Production AI Agents
CUSTOMER STORY | Driving $66M in revenue with predictive customer acquisition
Getting error in databricks agent response
Announcement | Automatic Upgrades: best practice features for your lakehouse tables
Labs
Databricks Certified Data Engineer Associate - Unable to Fetch My Certificate
Databricks Certified Data Engineer Associate - Download Certificate
Turning Databricks Incident Logs Into Diagrams People Actually Read
Data Quality Dashbaord - I have a requirement to build a custom designed quality dashboard in dbx
Understanding EXPLAIN FORMATTED in Databricks SQL
R plots are not rendering properly again
How do you keep a plain Delta copy of a streaming table? CLONE is not supported on streaming tables
Rethinking Database Storage: Why LTAP (Lakebase) is the Next Paradigm Shift
Materialized View backing pipeline retains old Unity Catalog after catalog rename
Issue: Lakeflow Connect Microsoft Teams Community Connector - No module named 'databricks.labs'
Introducing the Genie Hub: Ask Questions, Share Builds, and Master Conversational Analytics
comprehensive guide or usecases for data engineering hands-on
The Open-Weight Revolution: A Game Changer for Our LLM Cost Optimization Odyssey
How to Track the Latest Databricks Feature Names (Complete Rename History)
Enable Foundation Model APIs for AWS Marketplace trial workspace
Databricks Marketplace: IP Protection, Job Compute, and Secret Management
I am planning to publish an application through Databricks Marketplace. The application contains proprietary business logic and processing algorithms that must not be accessible to customers after installation. Customers should be able to use the application to process data in their own Databricks environment, but they should not be able to inspect, copy, or reuse the underlying implementation. Approaches I Have Tried I have considered the following approaches: Packaging the implementation as Python wheels or other artifacts. Hosting the implementation as private Python packages. Deploying the application source from a private Git repository. With the packaging approach, the underlying Python implementation may still be accessible or inspectable from notebooks, workspace files, cluster environments, package caches, or other customer-accessible locations. This does not meet my IP-protection requirements. With private package hosting, the package installation requires an access token or API key. Providing a reusable package-registry credential to an application running in the customer's workspace creates a risk that the credential could be extracted or reused to download the private package outside the intended application flow. I also understand that Databricks Marketplace reviewers may require the submitted application code to be human-readable during the review process. Therefore, I am looking for an architecture that protects the production implementation without relying only on code obfuscation. Current Understanding My understanding is that a Databricks App runs in an isolated, containerized runtime. This provides separation between the application runtime and normal customer workspace resources. However, my workload includes processing large datasets. I understand that Databricks App compute is primarily intended to run the application, API, or user interface, rather than perform large-scale Spark processing. Because of this, I am unsure whether Databricks App […truncated]
UI sends empty managed_identity_id, breaks storage credential creation with system-assigned Access
Learning Series | Advanced Machine Learning Operations
Certification not received even after 48 hours of passing the exam.
Ingesting data from a SharePoint Excel file with two worksheets in one pipeline
Solar EPC Services in India: Benefits, Companies & Guide
ISV Partnership
Unable to connect to Azure database from Databricks free edition
Databricks Cost for additional Workspace in Azure
🌟 Community Pulse: Your Weekly Roundup! July 13 – 19, 2026
NCC private endpoint cannot be added
INSERTS AND DELETES in a massive way for Lakeflow Spark Declarative Pipelines
Community Connectors confusion with `connector_spec.yaml` usage (or lack of usage)
Lakebase search
sdp-meta (dlt-meta) vs lakeflow_framework: when should we use which?
Handling Sensor Dropout in IoT Pipelines: A Quarantine Pattern with Lakeflow Declarative Pipelines
Lakeflow Community Connector Issue - ModuleNotFoundError: No module named 'databricks.labs'
Group manage permission using terraform
Branching databases like code: a CI/CD pattern for Lakebase at Glaspoort
Bug in ODBC Driver (SEGFAULT) - including full reproduction
Setup Before Trainning
LakeBridge in Databricks
I am doing data warehouse migration and at the last stage i.e reconcilliation . Now after running this command databricks labs lakebridge configure-reconcile --profile abhi it prompts me for selecting data source,report type,source catalog,target catalog , details for configuring reconcile metadata, and all got installed too but at the end got this error 14:38:46 ERROR [d.l.lakebridge.configure-reconcile] InvalidParameterValue: Only serverless compute is supported in the workspace. So my question is can we not do reconcilliation even after reconcilliation job got failed, if yes then how .
Multi-select dashboard parameter now resolves to NULL instead of empty array?
Learning Festival - No voucher received after completing required courses
Solution Accelerator Series | Social Determinants of Health
What is the compute for Lakeflow Connect SharePoint Connector
Lakeflow Connect & Community Connectors - 403 Errors + What is the compute?
Enforcing Minimum Cohort Size in a Databricks App with AI/BI Genie
Workload Identity Federation on ADO for account level provider
Snowflake OR Databricks
Advanced Learning Festival 15 June - 06 July 2026
New to the community!
Fraud detection in 27ms using Postgres for real-time feature lookups
Understanding the Databricks Genie Family: Which Genie Is Right for You?
Query Federation vs Lakeflow Connect in Databricks: When to Query and When to Ingest
Databricks Foreign Table
Indexes came to Lakehouse
Jobs and pipelines UI
😑Datafactory insists on Cluster compute BUT Databricks defaults to Serverless compute!
Generative AI training slides
Announcement | Foundational context: Cross-industry & function-specific accelerators for Lakebase
How to load PDFs incrementally in volume?
How Databricks Unity Catalog Business Semantics creates a governed layer for metrics
Is there a cost for Databricks Workflow notifications or are they free?
Request for 50% Certification Voucher Discount
Request for 50% Certification Voucher Discount
Databricks Partner Tiers Explained for the People Who Build the Delivery
registration code
Using Scala and Java UDFs in Unity Catalog
Week of Jul 13
68 questionsStop Translating Alteryx Boxes - A Lakebridge-assisted, test-driven migration to Azure Databricks
Unable to Login to Databricks free edition
Support Case #00957926 - No response regarding certification profile transfer
Databricks Generative AI Associate Certification
Building Plug-ins for AI/BI Dashboard
Databricks YAML files
Second Databricks Data Engineer Associate Exam Suspension – Looking for Advice
Databricks hits $188B valuation, extending its run as AI's favorite second act
Databricks Performance Optimization: What Changed, What Still Matters, and What Should Be Automated
Hello from a new member!
Full-text search
Run failed due to STORAGE_DOWNLOAD_FAILURE_SLOW → BOOTSTRAP_TIMEOUT
DLT pipeline cloning to another workspace.
CUSTOMER STORY | Unilever accelerates finance insights with Genie
AI-Powered Data Engineering with Lakeflow: Techniques for Modern Data Professionals | Virtual Event
Databricks is raising funding at $188B valuation
Building Multi Agent System with Supervisor Agent
Databricks raises $5B with a $188B valuation
AnalysisException: [UNRESOLVED_ROUTINE] Cannot resolve routine `=`
Small POCs Can Become Big Data + AI Solutions
Databricks Set to Hit $188B Valuation with New Investment from Coatue
--- top comments --- [jgalt212] raising funds at 100X revenue just like Elon. nicely done.
Lakebase CDF
Apache Spark 4.2 is officially here! Key architectural updates for AI-Native & Governed Platforms
Databricks App with Proprietary packages and code
The Boring Truth Behind a 60% Speedup
Matching Tables/Columns
My SQL looks something like this: select table_name, column_name from information_schema.columns where (column_name like '%C1%' or column_name like '%C2%' or column_name like '%C3%' or column_name like '%C4%' or column_name like '%C5%'); This will give me the matching table_name and column_name . But I also want to know which search string matched with column_name . The SQL syntax or way to query this is what I'm looking for. The result should look something like this: Table_Name Column_Name Search_String T1 T1_C1 C1 T2 T2_C2 C2 T2 T2_C3 C3 T3 T3_C4 C4
Not Every Pipeline Needs DLT — Here's How We Decided
Medallion Architecture in Practice: The Design Decisions Nobody Puts in the Diagram
Are enterprises moving from "Data Lakehouse" to "Agentic Lakehouse"?
What does "Agent-Ready Data Governance" actually mean in production?
New Agentic AI Ecosystem in Databricks
Governance for vector search across multi domain unstructured data
Show HN: Ingestr CDC – open-source CDC replication in Go
Hi all, this is Burak, one of the founders at Bruin. I created ingestr to make data ingestion easy. ingestr is a CLI tool that can ingest data from 130+ sources. I have shared it on HN after our Go rewrite as well, which made it the fastest ingestion tool in the space. However, ingestr had always been a batch tool. I am personally a big fan of batch workloads due to their simplicity and have built ingestr around that assumption as well; however, over time, the cracks started to show when we started working with larger orgs. Turns out there are some scenarios where CDC proves beneficial: - For legacy systems where it is not possible to introduce cursor columns due to technical, but mostly organizational, concerns, it becomes impractical to deploy batch pipelines. - For systems that do not have a way to reliably know the update timestamp, also due to legacy reasons. Think usecases where the columns are updated without the timestamp being updated. - For hard deletes. Even though I do believe there are ways to solve each of these, it ended up putting us in a disadvantage, and we decided to build the CDC connectors instead. ingestr CDC works in two modes now: batch (bad name, I know) and stream. The batch mode is reading the changelog entries from the last load until the starting timestamp, and the streaming mode keeps reading and landing them to the destination databases. ingestr has a few advantages compared to a more traditional debezium + kafka setup: - it's a standalone Go binary and does not require any additional infra. - it has very low resource consumption, ~100MB baseline. - it can run on your own computer during development, and can be converted into a streaming prod deployment when it is ready. - supports 20+ destinations already, primarily analytical platforms like snowflake, databricks, generic iceberg destinations, etc. i would love to hear any feedback on what we could do to make it easier for cdc workloads! https://github.com/bruin-data/ingestr --- top comments --- [mercutio93] Man I have had a lot of gripes with debezium... the setup alone in a kubernetes cluster with strimzi kafka connect... And dealing with the kafka connect errors and fixing replication slots... ehhh was painful. Does this tool make it any well easier. One problem I expirienced in production systems with staging doubles is that sometimes we do database restores from production onto staging and that makes the sinks and sources go haywire... does ingestr handle this well? Is support for kubernetes/helm anywhere in the pipeline
