Taming the Search: A Practical Way of Enforcing GDPR and CCPA in Large Datasets with Apache Spark
Description
In today’s data-driven economy, companies increasingly collect more user data as their valuable assets. By contrast, users have rightfully raised the concern of how to protect their data privacy. In response, there are data privacy laws to protect user’s privacy, among which, General Data Protection Regulation (GDPR) by European Union (EU) and California Consumer Privacy Act (CCPA) are two representative laws regulating business conduct in corresponding regions. Common requirements are to access or delete all records in all ever-collected data given a specific user’s search key(s) in a timely manner. The size of collected data and the volume of requests for search make enforcing GDPR and CCPA highly inefficient if not resourcefully infeasible. In this talk, we demonstrate our work for enforcing GDPR and CCPA in Adobe’s Experience Platform (AEP) by efficiently solve the search problem above. Specifically, we build Bloom Filters while saving data in Data Lake with minimal resource and maintenance overhead, which reduces nearly 10X searching time for a single search request. Furthermore, we build orchestrated microservices for splitting and scheduling extra-large search jobs into mul…
Description from YouTube. Full content on the video page.
Topics
More from Databricks
NewsTeach AI how your business actually runs
Model intelligence is no longer the bottleneck for enterprise AI adoption because modern frontier models easily handle complex reasoning tasks. Business value requires providing these models with specific organizational context and metadata about internal processes to create a competitive advantage.





