Skip to content
All topics

Compliance

Recent items mentioning Compliance across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.

22 recent items5 news14 videos3 community threads
What's happening in ComplianceAI synthesis · updated Aug 2026

Databricks paired Genie with Unity Catalog to hit 77% accuracy on agent tasks versus 56–72% for general coding agents, and Abacus Insights used this platform-native approach in a regulated environment to cut new-client onboarding time roughly in half 1. Separately, the incoming EU AI Act is reshaping healthcare and supply-chain AI deployments by classifying most clinical systems as high-risk, pushing radiology, diagnostic imaging, and clinical documentation use cases toward stricter regulatory scrutiny 23.

Written August 2026 from the 3 most recent items mentioning Compliance at that time. It refreshes when this topic next has enough new material. Click any [N] to jump to the source.

Reddit

Apache Spark ships with 28 known CVEs in production. In 2026. And nobody thinks this is a problem worth talking about publicly

This is not acceptable any longer in a modern developer world and community. I run security scans on our Spark deployment and came back with 28 vulnerabilities — high severity, public CVEs, all sitting in transitive dependencies like Netty and Apache Thrift. Nothing exotic. Netty published fixes for 22 of them in a single batch in June. Thrift fixed their issues in 0.23.0. The fixes exist. Spark just didn’t include them. What bothers me more than the CVEs themselves is the process — or the lack of one from Apaches side. In 2026, any serious DevSecOps pipeline is expected to have mandatory quality gates that block releases on high/critical CVEs in dependencies. This isn’t cutting-edge practice — it’s baseline hygiene. The full toolchain is free and mature There is no automated dependency CVE gate in Spark’s release pipeline. No Trivy. No Dependabot . No OWASP Dependency-Check. Nothing that would block a release because a bundled library has a known high-severity vulnerability. These are free tools. Adding one to a CI pipeline is an afternoon of work. It hasn’t been done. Problem Honest Assessment No automated dependency scanning Inexcusable in 2026. Free tools exist. One CI step. Decoupled release calendars Real coordination challenge, but solvable with Dependabot PRs that can be reviewed and merged quickly Volunteer PMC Doesn’t excuse Databricks, Google, Apple, and Amazon — all Spark committers with paid engineers — from contributing a security gate “Good enough” culture Actively harmful when Spark is used in AI/ML pipelines processing personal data under GDPR The July 2026 maintenance releases — 4.2.0, 4.1.3, 4.0.4, 3.5.9 — all shipped after Netty 4.1.135.Final was available . They didn’t include the bump. There was no public statement that the team was aware of the issue and working on it. It just shipped, vulnerable, into production systems everywhere. The “volunteer PMC” argument doesn’t land anymore. Databricks is worth somewhere around $62 billion. Google, Apple, Amazon, and Microsoft all have paid Spark committers . The resources to fix this governance gap before lunch exist. The will apparently doesn’t. The EU Cyber Resilience Act (CRA) — which came into force in 2024 and has mandatory compliance deadlines rolling in through 2027 — specifically targets software supply chain security, including transitive dependencies and SBOM (Software Bill of Materials) requirements. The ASF has acknowledged this directly, stating they need to prepare projects for “CRA and U.S. CISA guidance”. Spark ships into commercial products. This will become a compliance problem for a lot of companies very soon , and the fix is genuinely trivial from an engineering standpoint. The ASF did make progress in 2025 — launching “Apache Trusted Releases (ATR)” for distribution security and ratifying CycloneDX 1.7 for SBOM standards — but none of this yet translates to a blocking CVE gate on the release pipeline for projects like Spark . In the meantime: if you’re running Spark and using JFrog Xray or Trivy on your deployment, you can force-override the affected Netty and Thrift versions in your own build. It’s not clean but it works until the next maintenance release, expected sometime in Q4. Is anyone else tracking this or pushing upstream to get a CVE gate added to the build? submitted by /u/hrpedersen [link] [comments]

00hrpedersen1mo ago
HackerNews

Show HN: Matterbeam, a company-wide write-ahead log for your data

Hey HN. I'm Michael, founder of Matterbeam. Been chewing on the core ideas of it for over ten years, building toward it for three. demo: https://www.youtube.com/watch?v=YuhujARUmhA whitepaper: https://matterbeam.com/whitepaper Short version: companies build their data infra on point-to-point pipelines and one place to put all the data. Source A goes to warehouse B. Team C wants the same data shaped differently? Build another. Eventually a mess of brittle ETL nobody wants to touch. Matterbeam puts existing ideas together in a different way. Source data collected as immutable, time-ordered facts into a log. Destinations replay and transform those facts, from any point in time, into the target they need. One source, many uses. My last startup was acquired by Pluralsight in 2014. I ended up leading product architecture and data there for about five years. Working with really brilliant, product and data people that I would have said were doing everything _right_. Yet no one in the company was happy with data. It made me question if something more fundamental wasn't broken. A key inspiration came from Martin Kleppmann's 2015 talk "Turning the Database Inside Out." (https://www.youtube.com/watch?v=fU9hR3kiOK0) Most databases internally do something interesting: a write-ahead log (durable, append-only, time-ordered) as a source of truth, and derived structures are created (B-trees, indexes, materialized views) optimized to serve different read patterns. What if you took that pattern and blew it up to org scale? Your uses become materializations. Warehouse, RAG vector db, graph db, any new use created when needed with a late transform and a new emitter. A few comparisons: We aren't Kafka. Kafka is lower-level. My first attempt at this was at Pluralsight using Kafka as the log. It was crazy expensive and complicated to operate. For Matterbeam we built cloud-native: object storage gives durability, ephemeral compute avoids coordination, we don't need 100ms latency for most jobs. Allowed us to avoid a lot of Kafka's complexity. We aren't Fivetran. Fivetran is a managed pipeline. We're a utility. One customer replaced Fivetran when they brought us in. Saved them money, but that wasn't the goal, suddenly projects they estimated at five months started taking two days. A two-year migration compressed into months. Their PMs started asking to use Matterbeam for everything. We aren't a warehouse or lake. Snowflake and Databricks are great at what they're great at. The push to centralize all data in these systems was a mistake. We aim to be the layer underneath. Basically fulfill the original promise of the data lake: collect without a use case, materialize when you figure out what you need, in the shape and system you need. What's broken: This doesn't fit cleanly into "what does this replace" buckets. Most people agree data is broken but then lament "data is hard" or some form of "my team isn't doing it right." Nobody's actively looking to solve the deeper problem. Hard to find new customers even with glowing testimonials. Connector coverage. Fivetran has hundreds. We have way fewer in production. We're working on it, we're using AI, you can write your own pretty quickly. Still, if your stack needs fifty SaaS integrations on day one, we struggle. We're early. Handful of paid customers. Not large-enterprise-ready no SOC2, HIPAA etc yet. Also, conscious decision not to be open source. Long list of reasons, separate post. I'd love feedback on: How would you position or market this? It feels like category creation, which I know is hard. Does the mental model land, or is there a piece where you go WAT? If you've built CDC-into-warehouse, Kafka-plus-schema-registry, or rolled a data backbone, what's the part you'd have wanted an easier solution for? Blog, testimonials, marketing video on the site. I'll be watching the thread. Be brutal, I can take it (I think).

10mikepk4mo ago
RedditTutorial

Security Analysis Tool for Databricks (Deep-Dive w/ Arun, Principal Security Engineer @ Databricks)

SOC 2, HITRUST, and so many other security things! Security requires a multi-pronged approach. Databricks' Arun Pamulapati did a deep-dive into the Security Analysis Tool, a tool created by field engineers at Databricks to help you improve your organization's Databricks deployments security posture against threats. This is a very technical deep-dive and I hope you enjoy it! Link to repo: [https://github.com/databricks-industry-solutions/security-analysis-tool](https://github.com/databricks-industry-solutions/security-analysis-tool) Link to Databricks' security best practices: [https://www.databricks.com/trust/security-features/best-practices](https://www.databricks.com/trust/security-features/best-practices)

90JosueBogran5mo ago

Get Tuesday's version of this

Tracking Compliance? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.

Read past issues first