Skip to content
All videos
newsDatabricks·July 27, 2020

Running Apache Spark Jobs Using Kubernetes

Description

Apache Spark has introduced a powerful engine for distributed data processing, providing unmatched capabilities to handle petabytes of data across multiple servers. Its capabilities and performance unseated other technologies in the Hadoop world, but while Spark provides a lot of power, it also comes with a high maintenance cost, which is why we now see innovations to simplify the Spark infrastructure. Kubernetes on its right, offers a simplified way to manage infrastructure and applications. Kubernetes provides a practical approach to isolated workloads, limiting the use of resources, deploying on-demand and scaling as needed. Yaron Haviv will explain how to work with Kubernetes to build a single workflow with Spark based data preparation and ML tasks. Participants will learn how running Spark with Kubernetes enables users to unify analytics and data science on a single cloud-native architecture and eliminate the overhead of an extra big data cluster managed by different tools. About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.c

Description from YouTube. Full content on the video page.

More from Databricks