Skip to content
All topics

Apache Spark

Recent items mentioning Apache Spark across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.

60 recent items11 releases10 news10 videos29 community threads

What is Apache Spark?

Apache Spark is an open source, multi-language engine for data engineering, data science, and machine learning, running on anything from a single machine to a large cluster. On Databricks it's not an optional component: Spark is the technology powering every compute cluster and SQL warehouse, and each Databricks Runtime version ships a specific Spark version.

Spark matters on Databricks because almost everything you run ends up executing on it. PySpark DataFrames, SQL queries, Scala and R code, and Structured Streaming jobs for near real-time processing all share the same engine, so understanding how Spark distributes and optimizes work is the main lever for performance and cost on the platform.

The platform is now in the Spark 4 generation. Databricks Runtime 19 includes Apache Spark 4.2.0, while the long-term support lines carry Spark 4.1.0 (Runtime 18 LTS) and 4.0.0 (Runtime 17.3 LTS). Choosing a runtime is how you choose your Spark version, and newer runtimes are where new Spark capabilities show up first.

Which Spark version does Databricks run?

It depends on the Databricks Runtime you pick. Runtime 19 includes Apache Spark 4.2.0, Runtime 18 LTS includes 4.1.0, and Runtime 17.3 LTS includes 4.0.0. Older supported LTS lines such as 16.4 and 15.4 still carry Spark 3.5.

Do I need to install Spark to use Databricks?

No. Spark comes preinstalled and preconfigured as the engine behind every compute cluster and SQL warehouse. Picking a runtime version is the only Spark installation decision you make.

Is Apache Spark free?

Spark itself is open source software from the Apache Software Foundation and costs nothing to run on your own infrastructure. On Databricks you pay for the compute that executes your Spark workloads, not for the engine.

What's the difference between Spark and Databricks?

Spark is an open source engine you can run anywhere, from a laptop to a self-managed cluster. Databricks is a commercial platform in which Spark powers the compute layer, packaged into versioned runtimes with the rest of the platform built around it.

Sources: Apache Spark on Databricks, Databricks docs · Databricks Runtime release notes versions and compatibility, Databricks docs · Apache Spark project homepage

What's happening in Apache SparkAI synthesis · updated 1d ago

Databricks Runtime 18.3+ has made native IP functions generally available across PySpark, Scala, and SQL, leveraging Photon acceleration to execute IP CIDR joins up to 3.1x faster and 6.4x cheaper without custom UDFs 1. In the broader ecosystem, the release of delta-rs 1.0.0 resolved key Spark interoperability bugs affecting OPTIMIZE and MERGE operations alongside adding custom token credentials for Unity Catalog 4.

Generated daily from the 10 most recent items mentioning Apache Spark. Click any [N] to jump to the source.

Reddit

Data Quality in Microsoft Fabric – Native features vs. external tools (like Databricks Expectations)?

Hey everyone! I’ve been working a lot with data quality frameworks in Databricks (leveraging things like Delta Live Tables expectations or custom validation notebooks), and I'm curious about how the community is handling Data Quality in Microsoft Fabric. Does Fabric have a native, built-in feature for data quality/expectations (similar to what Databricks offers), or are most of you relying on external tools, libraries (like Great Expectations/Soda), or custom PySpark/SQL validation notebooks inside Data Engineering pipelines? How are you currently enforcing quality checks before data hits your Gold/Semantic layers in Fabric? Would love to hear your approaches and best practices! I read this documentation here: https://learn.microsoft.com/en-us/purview/unified-catalog-data-quality-fabric-lakehouse?WT.mc_id=510336 , but it doesn't say much. submitted by /u/SuperbNews2050 [link] [comments]

00SuperbNews20503d ago
Reddit

Time to Swap the Cookies for Jetfuel - New Dataset & New Databricks Genie Tutorial

Most of you know samples.bakehouse . Great for a first query. Perfect for a quick demo. But after years of cookie sales, it's a little overbaked. Time to swap the cookies for jet fuel. ✈️ Together with the OpenSky Network , I brought a full day of global air traffic to Databricks Marketplace: 696 million real ADS-B position reports, messy just like real life. Myself, I used Genie for the whole journey: EDA, data exploration, a Apache Spark Declarative Pipeline, and a Lakeflow Job. Then I went a step further and read the same Marketplace data with open-source tools only, using OpenSharing and pandas. The result is this hands-on tutorial: Marketplace + Unity Catalog: get the data as a governed table Genie Agents: find anomalies in plain English Genie Agents: explore and visualize with maps and charts Genie Code: a Spark Declarative Pipeline, bronze to gold, with data quality rules Genie Code: a Lakeflow Job with schedule, retries, and email alerts Databricks Apps: your coding agent, governed by Unity Catalog OpenSharing: the open-source client in VS Code with pandas Everything runs on Databricks Free Edition (free, no credit card). 📖 Tutorial: Databricks Genie for Data Engineers and Data Scientists 💻 GitHub: databricks/tmm/DSDE-Genie-Tutorial 🛫 Dataset: OpenSky Network full-day dataset on Marketplace What's the first thing you'd query in a day of global air traffic? P.S. For the record, we still love bakehouse! 🍪❤️ [Disclaimer: I'm one of the two people who baked it.] submitted by /u/CompetitiveBet8978 [link] [comments]

00CompetitiveBet89783d ago
Databricks CommunityDatabricks Academy Learners

Where to find Notebooks for course 'Advanced Techniques with Apache Spark Declarative Pipelines'

001w ago
Reddit

Data Engineering Project On Real Company | Real Data | Pyspark | Databricks | Production Ready

Hello Guys, I have come across many videos on YouTube related to Data Engineering, but I found very few that actually solve a real problem. Either the problem is made up, the data is unrealistic, or the whole project is built just for the sake of showing a project. I believe that if you really want to get hands on with Data Engineering, you first need to understand the **business, the data, and the actual solution the client demands.** So, that’s what we are going to do in this series. We are going to build a complete end to end Data Engineering project, starting from understanding the business problem and going all the way to building the actual solution. We will cover things like **Lakeflow Jobs, orchestration, data pipelines, and an actual company style workflow.** The idea is simple: not just build a project, but understand **why we are building it and how it would actually work in a company.** So without any further ado, let’s begin this series. submitted by /u/Formal_Ganache3620 [link] [comments]

00Formal_Ganache36201w ago
Reddit

How do you optimize slow PySpark jobs?

What do you usually check first—shuffle, partitioning, joins, Spark UI, file sizes, or something else? submitted by /u/Square-Designer7807 [link] [comments]

00Square-Designer78071w ago
Reddit

packaging data products

Hi, how do u handle data products, and chain of a lot of transformations, it can be very messy in the notebooks as mix of sqls, sqls python strings, python code, pyspark code, if/elses, etc... But again, is it really worth it to modularize it, make it callable, testable. I find it really hard to draw the line. How do u manage this in practice? submitted by /u/ptab0211 [link] [comments]

00ptab02111w ago
Reddit

Which Came First For You... Spark or Databricks?

Saw a post this morning claiming you shouldn't touch Databricks until you've already mastered Apache Spark. That advice is backwards. It treats learning like a 90s waterfall project and traps engineers in tutorial hell. The fundamentals are non-negotiable: - You need SQL and Python. - You need to understand distributed shuffles, partitions, and memory spills. - You need to understand Parquet structures, metadata footers, and row groups. - You need to understand Delta Lake transaction logs and ACID guarantees. If you don't know what happens under the hood, you will write brittle pipelines that burn cloud budgets. The false assumption is that you must learn these concepts in an abstract vacuum before touching a modern platform. You don't learn networking by memorizing RFC packet headers on a whiteboard for six months before opening Wireshark. You fire up the tool, capture packets, and inspect the traffic. Databricks, Fabric, or Snowflake is your network analyzer. It is the laboratory. Nobody in enterprise production is setting up bare-metal Spark clusters on a home lab just to learn how shuffles work. And running local[*] on a laptop masks real-world distributed realities: - Local filesystems use POSIX atomic renames. Cloud object stores (S3, ADLS) do not. You cannot truly grasp why Delta's _delta_log exists until you see it write commits against cloud storage. - A local JVM masks network serialization, executor skew, and cross-node shuffle overhead. The caveat: don't treat the platform as a black box. Don't just click "Run All" on serverless compute and assume you know Spark. Use the platform to inspect the engine: - Run .explain(True) and read the physical plan. - Open the Spark UI to watch tasks, stages, and memory spill. - Inspect the JSON commits inside _delta_log after a merge operation. - Compact small files and see how file pruning changes scan times. You don't master the theory first and then graduate to a platform. You master the fundamentals by using the platform as your workbench. Learn the platform. Learn the engine. Learn the storage layer. Do it concurrently, in the environment where production actually lives. submitted by /u/OkImprovement7010 [link] [comments]

00OkImprovement70101w ago
Reddit

PySpark vs Java/Scala for data pipelines

For those working with Spark in production: Do you mainly use PySpark, Scala, or Java? What do you use for batch vs streaming pipelines? Have you experienced significant performance issues with PySpark? If you've switched from PySpark to Java/Scala, did performance actually improve? At what point would you say it's worth learning Java/Scala instead of staying with Python? I'm wondering whether the performance difference is significant enough to justify switching to Java for production data pipelines. submitted by /u/almightysosa888 [link] [comments]

00almightysosa8882w ago
Databricks CommunityCommunity Articles

Learn Databricks Apache Spark Declarative Pipelines | From Associate to Professional

002w ago
Reddit

Read this if you use Streaming Tables in Lakeflow Spark Declarative Pipelines

🚀 We’re excited to announce that Lakeflow Spark Declarative Pipelines (SDP) now supports creating “vanilla” (i.e., non STREAMING) MANAGED TABLES and writing to them via one or more append flows , using the new CREATE TABLE ... FLOW ( SQL ) and create_table() (Python) APIs . What is this Beta? This Beta allows creating a managed table that is populated by append flows: CREATE TABLE ... FLOW (SQL) / create_table() + @append_flow (Python) create a managed table written by one or more flows. Fan multiple sources into one table — declare several flows targeting the same managed table. Full table surface works: partitioning, liquid clustering, expectations, row filters, table properties, and private (pipeline-local) tables. import_checkpoint on append_flow , which migrates an existing Structured Streaming workload into a pipeline without reprocessing the source — the flow imports the query's existing checkpoint and resumes from the last committed offset with state intact. Example (Python): from pyspark import pipelines as dp dp.create_table("combined") dp.append_flow(target="combined") def from_a(): return spark.readStream.table("source_a") u/dp.append_flow(target="combined") def from_b(): return spark.readStream.table("source_b") Example (SQL): CREATE TABLE events PARTITIONED BY (bucket) FLOW INSERT BY NAME SELECT id, bucket FROM STREAM read_files('abfss://my_path', format => 'json'); Where do we need help? We are in Beta, so there might be some rough edges. Please take this for a spin and share your feedback here . What’s next? Managed Tables support for other flow types (AutoCDC, Replace Using, and Replace Where) is coming soon! Learn more CREATE TABLE ... FLOW (SQL reference) — https://docs.databricks.com/aws/en/ldp/developer/ldp-sql-ref-create-table-flow create_table (Python reference) — https://docs.databricks.com/aws/en/ldp/developer/ldp-python-ref-create-table import_checkpoint on append_flow — https://docs.databricks.com/aws/en/ldp/developer/ldp-python-ref-append-flow Questions, feedback, or help: comment below or share feedback in the form: https://forms.gle/7bGP5FYN7P1Z4WP27 submitted by /u/SlightImagination250 [link] [comments]

00SlightImagination2503w ago
Databricks CommunityData Engineering

How to create monotonic function to incrementally add obj_id for datasource in pyspark

003w ago
Databricks CommunityTechnical Blog

Introducing Stream-Stream Join Support in Apache Spark Real-Time Mode

003w ago
Reddit

Acquiring/Processing from a MQ to a Delta

Has anyone tried acquiring data from a MQ at scale using apache spark on databricks cluster? I was trying to solve this problem at work but so far havn't seen an native lib or efficient ways to do this. The legacy system seems to be pulling data using a java based utility and wanted to if there are any imporvements or new patterns of access for spark based workflows. Any documentation or nudge is the right direction will be greatly appreciated. submitted by /u/raja3194 [link] [comments]

00raja31943w ago
Databricks CommunityData Engineering

Databricks vs Snowflake Pyspark Performance

003w ago
Reddit

Ingest image in databricks for powerbi ? a poc and any idea welcome

Spent some time this week on a POC that started from a business constraint: about 3,000 photos to ingest every week, sensitive enough that we can't just drop a shareable link in a dashboard, and they need to end up in Power BI where people already work, with row-level security. That combination rules out the easy answer. No public links, no loose files in blob storage floating around outside governance. The images had to live inside Delta on Databricks so Unity Catalog could handle access control, and Power BI had to be able to render them directly from the table. What the simple poc below does: Reads the images with Spark's binaryFile source (recursive lookup, glob filter on *.jpg) to pull path, modification time, size and raw bytes into one DataFrame. Encodes the binary content as a base64 data URL, so the image itself lives inside the row instead of behind a link. Then the actual blocker: Power BI caps text fields at roughly 32,766 characters, and a real photo's base64 string blows straight past that. So each string gets split into ~32,000-character segments and exploded into multiple rows, each tagged with its index and total length. On the Power BI side, a single DAX measure puts it back together in the right order before rendering: ​ Image_concat = IF( HASONEVALUE(images_in_delta[image_name]), CONCATENATEX( images_in_delta, images_in_delta[segment], , images_in_delta[split_index] ) ) Not elegant, but it's what gets a full-resolution image through a hard platform limit without touching the sensitivity requirement. The PySpark side, stripped to what matters — reading the images and doing the chunking: from pyspark.sql import functions as F from pyspark.sql import DataFrame #READ ALL THE IMAGES images_df = spark.read.format("binaryFile") \ .option("recursiveFileLookup", "true") \ .option("pathGlobFilter", "*.jpg") \ .load("/Volumes/main/image_ingest/image_sample") def add_base64url_from_image_binary(df: DataFrame, max_len: int = 32000) -> DataFrame: df_with_b64 = df.select( "*", F.concat(F.lit("data:image/jpg;base64,"), F.base64(F.col("content"))).alias("base64url") ) df_with_split_info = df_with_b64.select( "*", F.ceil(F.length(F.col("base64url")) / F.lit(max_len)).cast("int").alias("num_segments"), F.length(F.col("base64url")).alias("total_length") ) df_split = ( df_with_split_info .withColumn( "split_index", F.explode(F.sequence(F.lit(0), F.col("num_segments") - 1)) ) .select( "*", F.substring( F.col("base64url"), F.col("split_index") * max_len + 1, F.least(F.lit(max_len), F.col("total_length") - F.col("split_index") * max_len) ).alias("segment") ) .drop("base64url", "content") ) return df_split def add_image_name(df : DataFrame) -> DataFrame : return df.withColumn("image_name", F.regexp_replace(F.col("path"),".*/([^/]+)$", "$1")) df_images = add_base64url_from_image_binary(images_df) df_images = add_image_name(df_images) df_images.write.mode("overwrite").format("delta") \ .option("mergeSchema", "true") \ .saveAsTable("main.image_ingest.images_in_delta") Nothing here is exotic engineering — the interesting part was realizing early that the constraint wasn't really "how do we store images in Delta," it was "how do we get a sensitive image through Power BI's text field limit without ever exposing it outside the governed table." Once that was clear, the chunking workaround fell out naturally. At 3,000 images a week this holds up. If volume goes up meaningfully, I'd want to revisit whether inlining every image is still the right call versus resolving binary content on demand. Curious if others have hit the same Power BI ceiling with sensitive image data and landed on something cleaner than manual chunking. Have you any other idea than this ? submitted by /u/Data-space_men [link] [comments]

00Data-space_men1mo ago
Reddit

15+ years in ETL/Data, but relying heavily on AI (Copilot/Genie) lately. Am I still an engineer, or just a prompt validator?

Looking for a reality check from other data folks. I have 15+ years of experience in data (mostly Ab Initio and SQL), but moved to Databricks 2 years ago. I know data architecture and transformation logic well, but my raw Python coding skills are basic. My daily workflow usually looks like this: 1. I figure out the logic or root-cause the pipeline issue. 2. I use Copilot or Databricks Genie to generate the PySpark/Python code. 3. I review, test against edge cases, fix logic flaws, and validate the output. I’m great at step 3—I know how the data should behave. But because I rarely write code line-by-line from scratch anymore, I’ve been hit with huge imposter syndrome. It feels like I’m just a code reviewer for AI rather than a "real" engineer. Has anyone else from a traditional ETL background felt this shift in modern cloud stacks? Is this just the new reality of engineering, or am I letting my skills atrophy? submitted by /u/Terrible_Mud5318 [link] [comments]

00Terrible_Mud53181mo ago
Reddit

Apache Spark ships with 28 known CVEs in production. In 2026. And nobody thinks this is a problem worth talking about publicly

This is not acceptable any longer in a modern developer world and community. I run security scans on our Spark deployment and came back with 28 vulnerabilities — high severity, public CVEs, all sitting in transitive dependencies like Netty and Apache Thrift. Nothing exotic. Netty published fixes for 22 of them in a single batch in June. Thrift fixed their issues in 0.23.0. The fixes exist. Spark just didn’t include them. What bothers me more than the CVEs themselves is the process — or the lack of one from Apaches side. In 2026, any serious DevSecOps pipeline is expected to have mandatory quality gates that block releases on high/critical CVEs in dependencies. This isn’t cutting-edge practice — it’s baseline hygiene. The full toolchain is free and mature There is no automated dependency CVE gate in Spark’s release pipeline. No Trivy. No Dependabot . No OWASP Dependency-Check. Nothing that would block a release because a bundled library has a known high-severity vulnerability. These are free tools. Adding one to a CI pipeline is an afternoon of work. It hasn’t been done. Problem Honest Assessment No automated dependency scanning Inexcusable in 2026. Free tools exist. One CI step. Decoupled release calendars Real coordination challenge, but solvable with Dependabot PRs that can be reviewed and merged quickly Volunteer PMC Doesn’t excuse Databricks, Google, Apple, and Amazon — all Spark committers with paid engineers — from contributing a security gate “Good enough” culture Actively harmful when Spark is used in AI/ML pipelines processing personal data under GDPR The July 2026 maintenance releases — 4.2.0, 4.1.3, 4.0.4, 3.5.9 — all shipped after Netty 4.1.135.Final was available . They didn’t include the bump. There was no public statement that the team was aware of the issue and working on it. It just shipped, vulnerable, into production systems everywhere. The “volunteer PMC” argument doesn’t land anymore. Databricks is worth somewhere around $62 billion. Google, Apple, Amazon, and Microsoft all have paid Spark committers . The resources to fix this governance gap before lunch exist. The will apparently doesn’t. The EU Cyber Resilience Act (CRA) — which came into force in 2024 and has mandatory compliance deadlines rolling in through 2027 — specifically targets software supply chain security, including transitive dependencies and SBOM (Software Bill of Materials) requirements. The ASF has acknowledged this directly, stating they need to prepare projects for “CRA and U.S. CISA guidance”. Spark ships into commercial products. This will become a compliance problem for a lot of companies very soon , and the fix is genuinely trivial from an engineering standpoint. The ASF did make progress in 2025 — launching “Apache Trusted Releases (ATR)” for distribution security and ratifying CycloneDX 1.7 for SBOM standards — but none of this yet translates to a blocking CVE gate on the release pipeline for projects like Spark . In the meantime: if you’re running Spark and using JFrog Xray or Trivy on your deployment, you can force-override the affected Netty and Thrift versions in your own build. It’s not clean but it works until the next maintenance release, expected sometime in Q4. Is anyone else tracking this or pushing upstream to get a CVE gate added to the build? submitted by /u/hrpedersen [link] [comments]

00hrpedersen1mo ago
Reddit

Iceberg vs Deltalake (greenfield project with UC in 2026)

I saw the quarterly meeting and was quite shocked that Iceberg is prominently mentioned, (as much as Deltalake). Is it possible that both are getting the same amount of love from Databricks? Is anyone aware of the R&D effort on these formats, and can share conclusions from that? If I'm building a greenfield Unity Catalog, should I just flip a coin to decide what format to use? Here are the main concerns and priorities: Which one is better for OSS Apache Spark reads and writes Which one integrates with external software better (eg onelake shortcuts pointing from Fabric to Databricks UC). When Databricks is innovating within their own UC (eg. introducing new managed table functionality such as "MST Transactions"), which one of these formats are they likely to support first? Which are they likely to optimize better? Which format is more likely to remain 100% open source in the future (or as close to open source as required by customers who want portable blob data). Sorry if this appears to be a common question. I am a Databricks outsider. I am more familiar with Microsoft Fabric. Where that Fabric SaaS is concerned, you can be certain that Deltalake receives a LOT more promotion than Iceberg does. We rarely come across Iceberg, and it probably wouldn't appear in any marketing slide decks. We are likely to create a gold/presentation layer in UC soon. It will basically be created from scratch. It would be nice to know which of these parquet-based formats to pick, when presented with the choice. I understand there is lip-service given to both, and it claims that this choice "doesn't matter". But that doesn't necessarily take into account the potential integrations that are needed with external software (eg. for the benefit of exposing the same tables in onelake). Is one safer than the other? Is one of them a better choice for forward-looking purposes? submitted by /u/SmallAd3697 [link] [comments]

00SmallAd36971mo ago
Databricks CommunityCertifications

request for Databricks Certified Associate Developer for Apache Spark 3.0 free voucher

001mo ago
Databricks CommunityCommunity Articles

Keep a PySpark Test Flood Out of Your Coding Agent's Context Window

002mo ago
Databricks CommunityData Engineering

Apache Spark 4.2 is officially here! Key architectural updates for AI-Native & Governed Platforms

002mo ago
Databricks CommunityCommunity Articles

AI Technical Debt in Databricks: Why Generated Spark SQL and PySpark Still Need Metadata, Review

002mo ago
Databricks CommunityData Engineering

StatusCode.UNIMPLEMENTED error: DatabricksConnect library using AKS/PySpark to calling Spark cluster

003mo ago
Databricks CommunityData Engineeringanswered

How does Databricks handle registration and discovery of custom PySpark data sources in SDPs?

003mo ago
Databricks CommunityData Engineeringanswered

PySpark AnalysisException: Ambiguous reference to field t when parsing nested JSON

003mo ago
Databricks CommunityCommunity Articles

DataFlint on Databricks - the Open Source Spark UI Upgrade Apache Spark Has Needed for Years

003mo ago
Databricks CommunityTechnical Blog

Apache Spark’s Real-Time Mode Use Case Deep Dive: Gaming Sessionization

004mo ago
Databricks CommunityData Engineering

Apache Spark Masterclass (In-Person, Bengaluru) | 6 June

004mo ago
Databricks CommunityCommunity Articles

Converting stored procedures to PySpark

004mo ago

Get Tuesday's version of this

Tracking Apache Spark? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.

Read past issues first