Skip to content
brickster.ai
All videos
newsDatabricks·July 7, 2025

What’s New in Apache Spark™ 4.0?

Description

Join this session for a concise tour of Apache Spark™ 4.0’s most notable enhancements: SQL features: ANSI by default, scripting, SQL pipe syntax, SQL UDF, session variable, view schema evolution, etc. Data type: VARIANT type, string collation Python features: Python data source, plotting API, etc. Streaming improvements: State store data source, state store checkpoint v2, arbitrary state v2, etc. Spark Connect improvements: More API coverage, thin client, unified Scala interface, etc. Infrastructure: Better error message, structured logging, new Java/Scala version support, etc. Whether you’re a seasoned Spark user or new to the ecosystem, this talk will prepare you to leverage Spark 4.0’s latest innovations for modern data and AI pipelines. Talk By: Daniel Tenedorio, Sr. Staff Software Engineer, Databricks ; Wenchen Fan, Senior Staff Software Engineer, Databricks Here’s more to explore: Production ready data pipelines for analytics and AI: https://www.databricks.com/solutions/data-engineering The Big Book of Data Engineering: https://www.databricks.com/resources/ebook/big-book-data-engineering-2nd-edition See all the product announcements from Data + AI Summit: https://www

Description from YouTube. Full content on the video page.

More from Databricks