Skip to content
All videos
newsDatabricks·July 28, 2020

Everyday Probabilistic Data Structures for Humans

Description

Processing large amounts of data for analytical or business cases is a daily occurrence for Apache Spark users. Cost, Latency and Accuracy are 3 sides of a triangle a product owner has to trade off. When dealing with TBs of data a day and PBs of data overall, even small efficiencies have a major impact on the bottom line. This talk is going to talk about practical application of the following 4 data-structures that will help design an efficient large scale data pipeline while keeping costs at check. 1. Bloom Filters 2. Hyper Log Log 3. Count-Min Sketches 4. T-digests (Bonus) We will take the fictional example of an eCommerce company Rainforest Inc and try to answer the business questions with our PDT and Apache Spark and and not do any SQL for this. 1. Has User John seen an Ad for this product yet? 2. How many unique users bought Items A , B and C 3. Who are the top Sellers today? 4. Whats the 90th percentile of the cart Prices? (Bonus) We will dive into how each of these data structures are calculated for Rainforest Inc and see what operations and libraries will help us achieve our results. The session will simulate a TB of data in a notebook (streaming) and will have code sam

Description from YouTube. Full content on the video page.

More from Databricks