Skip to content
Apache SparkChangelog · September 2026

Databricks introduced on-demand state repartitioning for Apache Spark Structured Streaming queries.

Databricks introduced on-demand state repartitioning for Apache Spark Structured Streaming queries [12]. Running on Databricks Runtime 18 or above with the RocksDB state store provider, data engineers can now repartition state across an updated partition count simply by updating a configuration parameter and restarting the application [12]. This operational change eliminated the long-standing requirement to rebuild checkpoint directories or run costly migration paths when resizing long-running stateful pipelines [12].

In the wider open source ecosystem, the Delta-rs 1.0.0 release addressed critical Spark interoperability bugs that previously affected OPTIMIZE and MERGE operations, alongside adding support for column mapping and deletion vectors [3]. Community members also examined newly introduced stream-stream join capabilities within Spark real-time workloads [15], while others outlined best practices for deploying streaming tables using Spark declarative pipelines [13].

Community forums focused on runtime performance and practical PySpark pipeline patterns. Developers shared strategies for diagnosing slow PySpark workloads [6] and debated the trade-offs of PySpark compared to compiled Scala or Java pipelines [10]. Performance comparisons between Spark on Databricks and PySpark on Snowflake also prompted detailed analysis [17]. Finally, practitioners worked through implementation challenges such as creating monotonic object identifiers during incremental data ingestion [14] and finding resources for declarative pipeline certification courses [4, 11].

Everything cited

  1. [1]Data Quality in Microsoft Fabric – Native features vs. external tools (like Databricks Expectations)? community · 2026-09-29
  2. [2]Time to Swap the Cookies for Jetfuel - New Dataset & New Databricks Genie Tutorial community · 2026-09-29
  3. [3]rust-v1.0.0 release · 2026-09-28
  4. [4]Where to find Notebooks for course 'Advanced Techniques with Apache Spark Declarative Pipelines' community · 2026-09-25
  5. [5]Data Engineering Project On Real Company | Real Data | Pyspark | Databricks | Production Ready community · 2026-09-23
  6. [6]How do you optimize slow PySpark jobs? community · 2026-09-22
  7. [7]packaging data products community · 2026-09-22
  8. [8]Which Came First For You... Spark or Databricks? community · 2026-09-22
  9. [9]python-v1.6.4 release · 2026-09-21
  10. [10]PySpark vs Java/Scala for data pipelines community · 2026-09-18
  11. [11]Learn Databricks Apache Spark Declarative Pipelines | From Associate to Professional community · 2026-09-17
  12. [12]Announcing On-Demand State Repartitioning for Apache Spark™ Structured Streaming on Databricks news · 2026-09-14
  13. [13]Read this if you use Streaming Tables in Lakeflow Spark Declarative Pipelines community · 2026-09-11
  14. [14]How to create monotonic function to incrementally add obj_id for datasource in pyspark community · 2026-09-11
  15. [15]Introducing Stream-Stream Join Support in Apache Spark Real-Time Mode community · 2026-09-09
  16. [16]Acquiring/Processing from a MQ to a Delta community · 2026-09-06
  17. [17]Databricks vs Snowflake Pyspark Performance community · 2026-09-05
  18. [18]MLflow 3.16.0 release · 2026-09-04
  19. [19]138: Databricks Apps Explained | Part 2: Code Walkthrough video · 2026-09-03
  20. [20]Ingest image in databricks for powerbi ? a poc and any idea welcome community · 2026-09-01

A frozen monthly snapshot, generated from the brickster.ai archive and never rewritten. For the live view of this topic, see the Apache Spark hub. brickster.ai is an independent community project, not affiliated with Databricks, Inc.