Trending GitHub projects.
Repos tagged topic:databricks — open-source tools, integrations, and accelerators built on or around Databricks. Excludes the official repos already covered by Releases.
dbeaver
Free universal database tool and SQL client
redash
Make Your Company Data Driven. Connect to any data source, easily visualize, dashboard and share your data.
cube
📊 Cube Core is open-source semantic layer for AI, BI and embedded analytics
APIJSON
🏆 Real-Time no-code, powerful and secure ORM 🚀 providing APIs and Docs without coding by Backend, and Frontend(Client) can customize response JSONs 🏆 实时 零代码、全功能、强安全 ORM 库 🚀 后端接口和文档零代码,前端(客户端) 定制返回 JSON 的数据和结构
WrenAI
GenBI (Generative BI) for AI agents, an open-source, governed text-to-SQL through an open context layer that turns natural-language questions into trusted dashboards, charts, and SQL across 20+ data sources, such as BigQuery, Snowflake, PostgreSQL, ClickHouse, Amazon Redshift, Databricks and more.
sqlglot
Python SQL Parser and Transpiler
drawio-skill
Generate draw.io diagrams from natural language — 11 presets (UML, SysML/MBSE, BPMN, network, C4…), 39 tools: codebase/CI/infra-to-diagram, image→editable diagram, Databricks product icons, mind maps, build-up animation, exec-view compression, click-through runbooks, PR diff bot. Vision self-check, 10,000+ shapes. Exports PNG/SVG/PDF/JPG.
growthbook
Open Source Feature Flags, Experimentation, and Product Analytics
SynapseML
Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark
spark
.NET for Apache® Spark™ makes Apache Spark™ easily accessible to .NET developers.
ai-dev-kit
Databricks Toolkit for Coding Agents provided by Field Engineering
multiwoven
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
nao
👾 nao is an open source analytics agent. (1) Create context with nao-core cli, (2) deploy nao chat interface for everyone
zingg
Scalable master data management, identity resolution, entity resolution, and deduplication using ML
datacontract-cli
Enforce Data Contracts
altimate-code
Open-source agentic data engineering harness for dbt, SQL, and cloud warehouses. 100+ tools, 10 warehouses, AI-powered.
mlops-stacks
This repo provides a customizable stack for starting new ML projects on Databricks that follow production best-practices out of the box.
drawio-ai-kit
Teach your AI to draw correct, beautiful draw.io diagrams — declarative layout engine, ground-truth stencils, structural validator, vision self-check. AWS · Azure · GCP · Databricks · BPMN. Zero dependencies.
dataflare
Simple, easy-to-use database manager
Lynkr
Streamline your workflow with Lynkr, a CLI tool that acts as an HTTP proxy for efficient code interactions using Claude Code CLI.
sqlkit
Agentic SQL database manager for 40+ databases — PostgreSQL, MySQL, SQL Server, SQLite, DuckDB, ClickHouse, Oracle and more. Lightweight, privacy-first, AI-powered, cross-platform. Built with Tauri (Rust).
dbldatagen
Generate relevant synthetic data quickly for your projects. The Databricks Labs synthetic data generator (aka `dbldatagen`) may be used to generate large simulated / synthetic data sets for test, POCs, and other uses in Databricks environments including in Delta Live Tables pipelines
spark
Drop-in replacement for Apache Spark UI
dbx
🧱 Databricks CLI eXtensions - aka dbx is a CLI tool for development and advanced Databricks workflows management.
dqx
Databricks framework to validate Data Quality of pySpark DataFrames and Tables
databricks_bootcamp_2026
End-to-end Data Lakehouse project built on Databricks, following the Medallion Architecture (Bronze, Silver, Gold). Covers real-world data engineering and analytics workflows using Spark, PySpark, SQL, Delta Lake, and Unity Catalog. Designed for learning, portfolio building, and job interviews.
Dataflare
Fast. Simple. Database Manager.
terraform-databricks-examples
Examples of using Terraform to deploy Databricks resources
AI-Engineering-Lab
A free, self-paced 24-week AI engineering course: Python, machine learning, LLMs, RAG, fine-tuning, agents and MCP, Azure and Vertex and Bedrock, and Databricks. 43 runnable notebooks, one continuous case study. MIT licensed, no signup. By Zorost Intelligence AI Lab.
rocky
A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run — branches, replay, column-level lineage, compile-time contracts, per-model cost. Adapters: Databricks, Snowflake, BigQuery, DuckDB. Single static Rust binary. Apache 2.0.
lakehouse-engine
The Lakehouse Engine is a configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for several lakehouse algorithms, data flows and utilities for Data Products.
sdp-meta
Metadata driven Spark Declarative Pipelines framework for bronze/silver pipelines
dlt-meta
Metadata driven Spark Declarative Pipelines framework for bronze/silver pipelines
databricks-code-practice
Practice Databricks coding skills with hands-on exercises. Import into Databricks Free Edition, write code, run assertions, check pass/fail. Covers Delta Lake, Spark SQL, PySpark, Auto Loader, medallion architecture, window functions, and more.
databricks-sql-python
Databricks SQL Connector for Python
owox-data-marts
Open-Source Self-Service Analytics Platform
analytics-toolbox-core
A set of UDFs and Procedures to extend BigQuery, Snowflake, Redshift, Postgres and Databricks with Spatial Analytics capabilities
universql
Pushdown compute from Snowflake to DuckDB running on your infrastructure
stowage
Bloat-free, no BS cloud storage SDK.
databricks-apps-cookbook
Ready-to-use code snippets for building interactive Databricks Apps.
scalable-data-science
Scalable Data Science, course sets in big data Using Apache Spark over databricks and their mathematical, statistical and computational foundations using SageMath.
lakebridge
Accelerates migrations to Databricks by automating key migration activities
VariantSpark
machine learning for genomic variants
