Kafka
Recent items mentioning Kafka across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.
Plug & Play: Zerobus Ingest Now Supports Apache Kafka® Compatible APIs (Beta)
NewsLearn about Zerobus in 15 min!
Databricks Lakeflow Connect Zerobus Ingest is a high-performance, multi-cloud ingestion service that allows users to stream event data directly into their lakehouse without the cost and complexity of a traditional message bus. The video explains the architecture of Zerobus Ingest, announces upcoming API integrations for Kafka and MQTT, and demonstrates how to configure and run a Python client to write data directly into a Delta table.
NewsHow an Open, Scalable and Secure Data Platform is Powering Quick Commerce Swiggy's AI
Swiggy uses a Databricks-based lakehouse platform with Delta Lake to manage billions of daily events across 700+ cities, supporting real-time delivery predictions and resource optimization at 10-minute delivery speeds. The platform enables AI democratization through data agents and chatbots that provide automated customer support and business intelligence without requiring SQL expertise.
NewsNo More Fragile Pipelines: Kafka and Iceberg the Declarative Way
Kafka topics and Iceberg tables have incompatible data models that cause problems with schema evolution, consistency, and file fragmentation when building data pipelines. Confluent's Table Flow solution declaratively handles schema evolution, manages consistency through two-phase commits with segments, and performs automatic compaction to optimize query performance.
NewsStreaming Meets Governance: Building AI-Ready Tables With Confluent Tableflow and Unity Catalog
Confluent Tableflow automatically exposes Kafka streaming topics as Delta tables and syncs them to Databricks Unity Catalog, eliminating manual data pipeline work. A live demo shows how this enables immediate real-time analytics on stock market data streams without traditional ETL preprocessing steps.
NewsSponsored by: Confluent | Turn SAP Data into AI-Powered Insights with Databricks
Confluent streams SAP operational data, enriches it with real-time context, and feeds it to Databricks for machine learning and analytics. The platform enables bidirectional data flow between operational and analytical systems, allowing AI agents to access both historical data and fresh signals while returning insights back to production applications.
NewsCreating a Custom PySpark Stream Reader with PySpark 4.0
The video demonstrates how to create a custom PySpark 4.0 streaming data source by implementing Python classes that inherit from the new data source interfaces. It walks through building a custom stream reader to ingest data from an unsupported messaging system like Apache Active MQ directly into Delta tables.
NewsPractical Pipelines: A Houseplant Alerting System with ksqlDB
This video demonstrates how to build a real-time houseplant monitoring system using Apache Kafka and a Raspberry Pi. It teaches how to use ksqlDB for stateful stream processing, configure Kafka topics, and integrate hardware sensors with external messaging alerts.
NewsUnlocking Near Real Time Data Replication with CDC, Apache Spark™ Streaming, and Delta Lake
DoorDash implemented an open-source near real-time data ingestion framework called Pepto using Apache Kafka, Apache Spark Streaming, and Delta Lake. The system captures change data feeds from various transactional databases and performs micro-batched upserts to achieve a P99 data latency of 7 to 30 minutes.
Get Tuesday's version of this
Tracking Kafka? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.











