Snowflake
Recent items mentioning Snowflake across the Databricks ecosystem — releases, news, videos, and community Q&A. Updated hourly.
dbt Labs recognized Snowflake as a 2026 Partner of the Year alongside key ecosystem integrators like phData and 66degrees 3. On the integration front, practitioners are standardizing secure cross-platform access by connecting Databricks to Snowflake via Workload Identity Federation 1.
Generated daily from the 4 most recent items mentioning Snowflake. Click any [N] to jump to the source.
Connecting to Snowflake via Workload Identity Federation
Which Came First For You... Spark or Databricks?
Saw a post this morning claiming you shouldn't touch Databricks until you've already mastered Apache Spark. That advice is backwards. It treats learning like a 90s waterfall project and traps engineers in tutorial hell. The fundamentals are non-negotiable: - You need SQL and Python. - You need to understand distributed shuffles, partitions, and memory spills. - You need to understand Parquet structures, metadata footers, and row groups. - You need to understand Delta Lake transaction logs and ACID guarantees. If you don't know what happens under the hood, you will write brittle pipelines that burn cloud budgets. The false assumption is that you must learn these concepts in an abstract vacuum before touching a modern platform. You don't learn networking by memorizing RFC packet headers on a whiteboard for six months before opening Wireshark. You fire up the tool, capture packets, and inspect the traffic. Databricks, Fabric, or Snowflake is your network analyzer. It is the laboratory. Nobody in enterprise production is setting up bare-metal Spark clusters on a home lab just to learn how shuffles work. And running local[*] on a laptop masks real-world distributed realities: - Local filesystems use POSIX atomic renames. Cloud object stores (S3, ADLS) do not. You cannot truly grasp why Delta's _delta_log exists until you see it write commits against cloud storage. - A local JVM masks network serialization, executor skew, and cross-node shuffle overhead. The caveat: don't treat the platform as a black box. Don't just click "Run All" on serverless compute and assume you know Spark. Use the platform to inspect the engine: - Run .explain(True) and read the physical plan. - Open the Spark UI to watch tasks, stages, and memory spill. - Inspect the JSON commits inside _delta_log after a merge operation. - Compact small files and see how file pruning changes scan times. You don't master the theory first and then graduate to a platform. You master the fundamentals by using the platform as your workbench. Learn the platform. Learn the engine. Learn the storage layer. Do it concurrently, in the environment where production actually lives. submitted by /u/OkImprovement7010 [link] [comments]
Celebrating the 2026 dbt partner of the year winners
dbt Labs named its 2026 Partner of the Year winners: phData, Snowflake, Cívica, Datum Studio, and 66degrees. For Databricks practitioners tracking the broader analytics-engineering ecosystem, it's a signal of which implementation and platform partners dbt Labs sees as driving the most impact this year.
Show HN: Like LeetCode, but multi-file and multi-step
On LeetCode you write one function, it passes, and you're done. Most AI labs such as Anthropic and OpenAI, as well as some big tech companies such as Meta, Snowflake and Databricks, don't stop in one function and instead ask multi-step questions that are closer to real-world problems. Say you're building a key-value store. Step 1 is get, set, delete, count. Once your tests pass, Step 2 unlocks: expiring keys. Then nested transactions. Then atomic batches. I built this because I couldn't find a good place to practice for multi-step, multi-file questions. The questions are based on ones people have reported from those loops. Check it out at https://loopprep.dev. There are 24 questions, 3 free, and no signup to try them. Happy to answer questions. --- top comments --- [MiroslavPokorny] Writing code in a browser like this or leet code is wrong on so many levels its not funny.
Need some advice on Snowflake vs Databricks
submitted by /u/Ok_Independent_343 [link] [comments]
Snowflake’s AI generated slop blog
submitted by /u/Useful-Swordfish4946 [link] [comments]
Databricks vs Snowflake comparison
Are there any unbiased comparisons between these two popular platforms? Seen a lot but most of them are biased views, based on experience and commercial motives. submitted by /u/medici2022 [link] [comments]
Databricks vs Snowflake Pyspark Performance
Data Engineers, what does your actual day-to-day work look like? And what should I learn next?
I’m currently trying to transition deeper into Data Engineering and would really appreciate some perspective from people who are already working in the field. I have 1.3 yrs experience as a Junior Python Developer. What I want to do is slowly transform into a Data Engineer. How would you suggest my choice? Basically what I do is make web scraping scripts to get the data from web and give the data in excels. Our company is currently not using git or CI/CD or anything like that. The problem I’m running into is that when I look at Data Engineering jobs on Naukri, LinkedIn, etc., the requirements seem endless. One job asks for Python, SQL, Airflow and AWS; another wants Spark, Kafka and Databricks; another wants Snowflake, dbt, Terraform, Kubernetes, CI/CD, etc. It becomes difficult to understand what I should actually prioritize. So I’d like to hear from people who are actually working as Data Engineers . What does your day-to-day work look like? What kind of problems do you solve, what technologies do you use regularly, and which skills have turned out to be genuinely important in your job? More importantly, based on my current experience, what would you suggest I improve or learn next to become a stronger candidate for Data Engineering roles? Are there any gaps that you think I should focus on, or technologies/concepts that are worth learning through projects rather than just studying theoretically? I’m not really looking for a generic “learn SQL → Python → Spark → AWS” roadmap. I’m more interested in understanding the reality of the job and getting advice from people who have actually gone through the transition. If you’re a Data Engineer with 1–5+ years of experience , I’d especially appreciate your perspective. Even a short description of what you work on and what you wish you had learned earlier would be extremely helpful. Thanks in advance! submitted by /u/Stunning-Space8032 [link] [comments]
Lakehouse Federation (Snowflake) — large query results fail to download from internal stage
Convert proprietary code to open ANSI SQL with Genie Code
Now in Beta, Genie Code powers an agentic converter that launches swarms of parallel agents to translate proprietary SQL into open ANSI SQL, iteratively validating both syntax and semantic intent. It covers T-SQL, Snowflake, Redshift, Oracle, BigQuery, and Teradata sources, with migration projects in the Databricks workspace to track progress, visualize lineage, and identify objects that need to move together.
Retiring the dbt Snowflake Native App
dbt Labs is retiring the dbt Snowflake Native App from the Snowflake Marketplace, placing it into maintenance mode in July 2026 before full removal in November 2026. The sunset only affects teams with an active Native App installation and introduces no changes to dbt platform, dbt Core, the dbt Semantic Layer, or standard data workflows.
Snowflake OR Databricks
Show HN: Ingestr CDC – open-source CDC replication in Go
Hi all, this is Burak, one of the founders at Bruin. I created ingestr to make data ingestion easy. ingestr is a CLI tool that can ingest data from 130+ sources. I have shared it on HN after our Go rewrite as well, which made it the fastest ingestion tool in the space. However, ingestr had always been a batch tool. I am personally a big fan of batch workloads due to their simplicity and have built ingestr around that assumption as well; however, over time, the cracks started to show when we started working with larger orgs. Turns out there are some scenarios where CDC proves beneficial: - For legacy systems where it is not possible to introduce cursor columns due to technical, but mostly organizational, concerns, it becomes impractical to deploy batch pipelines. - For systems that do not have a way to reliably know the update timestamp, also due to legacy reasons. Think usecases where the columns are updated without the timestamp being updated. - For hard deletes. Even though I do believe there are ways to solve each of these, it ended up putting us in a disadvantage, and we decided to build the CDC connectors instead. ingestr CDC works in two modes now: batch (bad name, I know) and stream. The batch mode is reading the changelog entries from the last load until the starting timestamp, and the streaming mode keeps reading and landing them to the destination databases. ingestr has a few advantages compared to a more traditional debezium + kafka setup: - it's a standalone Go binary and does not require any additional infra. - it has very low resource consumption, ~100MB baseline. - it can run on your own computer during development, and can be converted into a streaming prod deployment when it is ready. - supports 20+ destinations already, primarily analytical platforms like snowflake, databricks, generic iceberg destinations, etc. i would love to hear any feedback on what we could do to make it easier for cdc workloads! https://github.com/bruin-data/ingestr --- top comments --- [mercutio93] Man I have had a lot of gripes with debezium... the setup alone in a kubernetes cluster with strimzi kafka connect... And dealing with the kafka connect errors and fixing replication slots... ehhh was painful. Does this tool make it any well easier. One problem I expirienced in production systems with staging doubles is that sometimes we do database restores from production onto staging and that makes the sinks and sources go haywire... does ingestr handle this well? Is support for kubernetes/helm anywhere in the pipeline
Show HN: Connect DuckDB to any database that has an ADBC driver
Hi everyone, I'm a PhD student in databases at CMU. Over the past few months, I've been interning at Columnar and building a community extension for DuckDB that lets you query Snowflake, Databricks, BigQuery, PostgreSQL, MySQL, and any other system with an ADBC (Arrow Database Connectivity) driver. The extension supports querying ADBC databases directly through a read_adbc table function. It also supports using ATTACH to connect to an ADBC database and then running SELECT, INSERT, COPY, and CTAS statements as if the database were local to DuckDB. You can install it from DuckDB by running: INSTALL adbc FROM community; LOAD adbc; The ADBC extension is open source and available under the Apache-2.0 License. We plan to contribute this extension to the Arrow project so it can become an official ADBC library. Give it a try, and leave feedback if you have ideas about how to make it better. Leave an issue on the GitHub repo if you hit any rough edges! --- top comments --- [eveerett] Super cool project! Curious your thoughts on LakeSail. It can run single node on your laptop similar to DuckDB all the way to a cluster like Spark (it adopts the Spark interface, but built entirely in Rust without the JVM). Might be an interesting look for ya. [StarlaAtNight] Haven’t tried it yet, but awesome work! This seems huge. I try to use DuckDB whenever feasible, and I also tend to work with Snowflake, Databricks, and BigQuery (as well as Fivetran which uses DuckDB for local dev/debug mode). Excited to have ADBC-enabled connections w/ DuckDB, and have a feeling this is going to start coming up as a use case a lot more
Show HN: Dex – Cost-aware analytics engineering skills for agents
Hi I’m Marco, co-founder of Exmergo. Me and my team created Dex to help Analytics Engineers do real work with Claude Code (and any other coding agent). We’ve found that Data and Analytics teams are stuck between a rock and a hard place cost-wise: - On one side, they are using some of the most expensive consumption-billed software on the planet (Snowflake, Databricks etc.). - On the other side, Anthropic and Open AI want you to tokenmax (and with data analytics it’s very easy to burn your context window). So we created Dex, our open source skills plugin (Apache-2.0), to solve both of these problems: - Dex forces the agent to use a cost guard when performing exploration queries and transformations. - Dex uses a set of tooling that makes it hard to get burned when transforming data. It achieves these things with tight control scripts that the SKILL.md files are pointed towards when using /dex:explore, /dex:transform and /dex:maintain. Bonus: Dex makes agents really good data exploration, building sql and dbt models and detecting drift. 76% performance on ade-bench with Claude Sonnet 5 (and, per our measures, 2.5x cheaper than Fable 5). Install on any agent with this command in your terminal: npx skills add exmergo/dex Install on Claude Code running these commands (separately): /plugin marketplace add exmergo/exmergo-agent-plugins /plugin install dex@exmergo If you want to see more visual examples you can go through the README or browse here: https://www.exmergo.com/dex Let me know if this helps your analytics workflow and makes your agents more cost-aware.
Ingest data from snowflake to databricks
dbt Labs Named Snowflake Data Integration Product Partner of the Year
dbt Labs was named Snowflake Data Integration Product Partner of the Year. This post details dbt Labs' two Snowflake Partner honors, including the CoCo Adoption Award.
What we announced at Snowflake Summit and why it matters
dbt Labs has open-sourced its Rust-powered Fusion runtime as dbt Core v2.0, bringing up to 10x faster parse times and day-one support for the Databricks adapter. The update also launches dbt State in preview, enabling state-aware orchestration that skips unchanged models to cut compute costs by an average of 30 percent.
Snowflake to Databricks Migration in 12 weeks and cut cost per run by ~77%. AMA.
Lovelytics wrapped up a Snowflake-to-Databricks migration; 847 DBT models, 35 Info Mart tables, \~77% lower cost per run on a 2XL warehouse. **TL;DR What helped:** * Treated the migration as engineering, not translation. Each dbt model was tested in isolation, not just row counts vs Snowflake. * Routing macro to resolve cross-layer references at runtime, so the same codebase could read from Snowflake, federated Snowflake, and Unity Catalog without forking logic. * Dual model trees in one repo, which let the migration stay in lockstep with live Snowflake changes. * Script-generated wave selectors enabled parallel builds while preserving dependency order. * Used reference-slice validation subsets vs. waiting on full mart refreshes. **TL;DR Cost reduction:** * Reworked joins to use narrow staging dimensions instead of wide marts where possible. * Added incremental predicates to reduce MERGE target scans. * Split wide models into parallel sub-models where the dependency graph allowed it. * Copied static reference data into Delta instead of repeatedly reading it through federation. * Loaded static copies into Delta rather than reading via federation (predicate pushdown is poor). Happy to go into the gotchas: HASH() not being portable, Snowflake MERGE tolerating duplicate keys that Delta doesn't, NULL ordering, and timestamp handling. AMA [Full Blog Post](https://community.databricks.com/t5/technical-blog/partner-blog-847-models-12-weeks-77-less-inside-r1-s-snowflake/ba-p/157284)
Sqlit – A lazygit-style TUI for SQL databases
sqlit is a lazygit-style TUI for SQL. Connect and query your database from the terminal in seconds, no config files or documentation to read first. It supports all the major databases: SQL Server, PostgreSQL, MySQL, SQLite, MariaDB, FirebirdSQL, Oracle, DuckDB, CockroachDB, ClickHouse, Snowflake, Databricks, Supabase, Cloudflare D1, Turso, Athena, BigQuery, Spanner, Redshift, IBM Db2, SAP HANA, Teradata, Trino, Presto, Apache Flight SQL, Apache Impala, SurrealDB, and osquery. A few things that come built in: - Keyboard focus: Context aware keybindings always visible - Docker integration: auto-detects and connect to running database containers - Vim-style query editor with customizable keybindings. - Fuzzy filter in results window. - SSH tunnels, OS-keyring credential storage, password manager integration. - Autocomplete for tables, columns, and procedures. - Cloud CLI integration (browse external DBs via Azure / AWS / GCP CLIs). - Themes (Rose Pine, Tokyo Night, Nord, Gruvbox). Install: `pipx install sqlit-tui` (also works with `uv tool install` and `pip`). Built with Python and Textual. First shared here in December (https://news.ycombinator.com/item?id=46276002) - a lot has shipped since. Repo: https://github.com/Maxteabag/sqlit Feedback welcome, especially on what's still missing for daily-driver use. My goal is to make an aesthetic tool that makes it easy and enjoyable to connect and query data, and do that one thing only, really well. --- top comments --- [cesarmanzocode] Very nice UX direction. One thing I’ve consistently missed in terminal DB tools is schema-aware query history/search across multiple databases. Feels like that would fit this aesthetic really well. [qsera] Looks amazing. But before I can commit to trying it, I would like to know how much LLM involvement was there. [teejmya] Why is this in Ask HN? Shouldn't it be in Show HN?
Why should I use Databricks Over Snowflake ?
Where do you Qlik/Talend with the likes of Databricks?
Hi Folks, Long time Qlik Analytics user and since Talend merger have been involved in data engineering side of things as well. While I already know the major differences and benefits Databricks/Snowflake and new age BI solutions like Sigma and Omni provides over other vendors. Wondering where does the market see vendors like Qlik ?
[PARTNER BLOG] 847 Models, 12 Weeks, 77% Less: Inside R1's Snowflake-to-Databricks Migration
Seeking Volunteers with Lakehouse, Fabric, Databricks, or Snowflake Experience
[BLOG + video] Snowflake and Databricks benchmarks
We put Snowflake and Databricks head-to-head across 5 scenarios. 𝗦𝗻𝗼𝘄𝗳𝗹𝗮𝗸𝗲 𝘄𝗼𝗻 𝟰 𝗼𝘂𝘁 𝗼𝗳 𝟱 𝘀𝗰𝗲𝗻𝗮𝗿𝗶𝗼𝘀: \- Sequential queries: 34% faster, 17% cheaper (at $2/credit) \- Concurrent queries: 38% faster, 39% cheaper \- Cold start: 54% faster (Databricks startup time: \~7 sec. Snowflake: sub-second. Every. Single. Time.) \- DML (delete + insert): 59% faster, 32% cheaper thanks to elite query pruning that treated 6B rows like 6M 𝗗𝗮𝘁𝗮𝗯𝗿𝗶𝗰𝗸𝘀 𝗰𝗹𝗮𝗶𝗺𝗲𝗱 𝘁𝗵𝗲 𝗼𝗻𝗲 𝘁𝗵𝗮𝘁 𝗺𝗮𝘁𝘁𝗲𝗿𝘀 𝗳𝗼𝗿 𝗱𝗮𝘁𝗮 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝘀: CTAS (Create Table As Select): 58% faster, 71% cheaper when writing billions of rows across multiple table shapes If your workload is heavy on dbt materializations, large table builds, or data pipeline writes, Databricks has a real edge here. If your workload is analysts running queries, dashboards, and incremental refreshes, Snowflake Standard looks compelling, even vs. Databricks Enterprise pricing. \- Read the full methodology and results: [https://select.dev/posts/snowflake-vs-databricks-showdown](https://select.dev/posts/snowflake-vs-databricks-showdown) \- Take a look at the repo: [https://github.com/get-select/snowflake-databricks-benchmark](https://github.com/get-select/snowflake-databricks-benchmark)
Show HN: Matterbeam, a company-wide write-ahead log for your data
Hey HN. I'm Michael, founder of Matterbeam. Been chewing on the core ideas of it for over ten years, building toward it for three. demo: https://www.youtube.com/watch?v=YuhujARUmhA whitepaper: https://matterbeam.com/whitepaper Short version: companies build their data infra on point-to-point pipelines and one place to put all the data. Source A goes to warehouse B. Team C wants the same data shaped differently? Build another. Eventually a mess of brittle ETL nobody wants to touch. Matterbeam puts existing ideas together in a different way. Source data collected as immutable, time-ordered facts into a log. Destinations replay and transform those facts, from any point in time, into the target they need. One source, many uses. My last startup was acquired by Pluralsight in 2014. I ended up leading product architecture and data there for about five years. Working with really brilliant, product and data people that I would have said were doing everything _right_. Yet no one in the company was happy with data. It made me question if something more fundamental wasn't broken. A key inspiration came from Martin Kleppmann's 2015 talk "Turning the Database Inside Out." (https://www.youtube.com/watch?v=fU9hR3kiOK0) Most databases internally do something interesting: a write-ahead log (durable, append-only, time-ordered) as a source of truth, and derived structures are created (B-trees, indexes, materialized views) optimized to serve different read patterns. What if you took that pattern and blew it up to org scale? Your uses become materializations. Warehouse, RAG vector db, graph db, any new use created when needed with a late transform and a new emitter. A few comparisons: We aren't Kafka. Kafka is lower-level. My first attempt at this was at Pluralsight using Kafka as the log. It was crazy expensive and complicated to operate. For Matterbeam we built cloud-native: object storage gives durability, ephemeral compute avoids coordination, we don't need 100ms latency for most jobs. Allowed us to avoid a lot of Kafka's complexity. We aren't Fivetran. Fivetran is a managed pipeline. We're a utility. One customer replaced Fivetran when they brought us in. Saved them money, but that wasn't the goal, suddenly projects they estimated at five months started taking two days. A two-year migration compressed into months. Their PMs started asking to use Matterbeam for everything. We aren't a warehouse or lake. Snowflake and Databricks are great at what they're great at. The push to centralize all data in these systems was a mistake. We aim to be the layer underneath. Basically fulfill the original promise of the data lake: collect without a use case, materialize when you figure out what you need, in the shape and system you need. What's broken: This doesn't fit cleanly into "what does this replace" buckets. Most people agree data is broken but then lament "data is hard" or some form of "my team isn't doing it right." Nobody's actively looking to solve the deeper problem. Hard to find new customers even with glowing testimonials. Connector coverage. Fivetran has hundreds. We have way fewer in production. We're working on it, we're using AI, you can write your own pretty quickly. Still, if your stack needs fifty SaaS integrations on day one, we struggle. We're early. Handful of paid customers. Not large-enterprise-ready no SOC2, HIPAA etc yet. Also, conscious decision not to be open source. Long list of reasons, separate post. I'd love feedback on: How would you position or market this? It feels like category creation, which I know is hard. Does the mental model land, or is there a piece where you go WAT? If you've built CDC-into-warehouse, Kafka-plus-schema-registry, or rolled a data backbone, what's the part you'd have wanted an easier solution for? Blog, testimonials, marketing video on the site. I'll be watching the thread. Be brutal, I can take it (I think).
Is Fabric just “good enough,” or does Databricks still win?
I’ve been at a few Microsoft centric events lately hearing this a lot lately: *“What’s the difference between Fabric, Databricks, and Snowflake and when would you choose each?”* I'm curious what tips the scale for you and what your one-line answer would be? * Team skillsets? * Data scale/complexity? * Cost control? * Governance? * AI/ML needs?
Show HN: Mljar Studio – local AI data analyst that saves analysis as notebooks
Hi HN, I’ve been working on mljar-supervised (open-source AutoML for tabular data) for a few years. Recently I built a desktop app around it called MLJAR Studio. The idea is simple: you talk to your data in natural language, the AI generates Python code, executes it locally, and the whole conversation becomes a reproducible notebook (*.ipynb file). So instead of just chatting with data, you end up with something you can inspect, modify, and rerun. What MLJAR Studio does: - Sets up a local Python environment automatically, runs on Mac, Windows, and Linux - Installs missing packages during the conversation - Built-in AutoML for tabular data (classification, regression, multiclass) - Works with standard Python libraries (pandas, matplotlib, etc.) - Works with any data file: CSV, Excel, Stata, Parquet ... - Connects to PostgreSQL, MySQL, SQL Server, Snowflake, Databricks, and Supabase. For AI: use Ollama locally (zero data egress), bring your own OpenAI key, or use MLJAR AI add-on. I built this because I wanted something between Jupyter Notebook (flexible but manual) and AI tools that generate code but don’t preserve the workflow. Most tools I tried either hide too much or don’t give reproducible results and are cloud based Demos: - 60-second demo: https://youtu.be/BjxpZYRiY4c - Full 3-minute analysis: https://youtu.be/1DHMMxaNJxI Pricing is $199 one-time, with a 7-day trial. Curious if this is useful for others doing real data work, or if I’m solving my own problem here. Happy to answer questions. --- top comments --- [MSaiRam10] Notebooks as the output format is funny because notebooks are famously bad for reproducibility. Out of order execution, hidden state, etc. You're solving "chat isn't reproducible" with a format that also isn't really [hasyimibhar] How does this compare to open source Deepnote[0]? We use the cloud version (BYOC) at my previous company to replace self-hosted Jupyter notebooks, and it's pretty great. [0] https://github.com/deepnote/deepnote [2ndorderthought] This is one of those product areas I would call high-risk without a human in the loop. So I am glad you kept a person in the loop. It's really easy to lose tons of money making decisions based on bad statistics or models. Anyone remember how much money zillow lost because of automatic time series models? I do have concerns about the workflow. Data people aren't usually the best programmers. Models hallucinate and make mistakes sometimes subtle sometimes not. Can you think of a way to prevent data scientists from having to be expert code reviewers? I feel like taking away the code gives them the chance to find and fix mistakes in their reasoning but I have no evidence for that. [amirathi] Really cool. If somebody doesn't want to adopt a new platform, take a look at open source Jupyter MCP Server[1]. Once integrated with Claude, it can execute code on the live notebook kernel. I just let Claude write notebooks, run top to bottom, debug & fix errors & only ping me when everything is working. [1] https://github.com/datalayer/jupyter-mcp-server [trymamboapp] "AI saves analysis as notebooks" is fighting the wrong fight ig. The reproducibility issue with notebooks isn't the format. it's out-of-order cell execution and silent kernel state llm generation makes that worse: the model has no memory of what state existed when it wrote cell 7, and neither does the user.
Show HN: Rocky – Rust SQL engine with branches, replay, column lineage
Hi HN, I'm Hugo. I've been building Rocky over the past month, shipping fast in the open. The binary is on GitHub Releases, `dagster-rocky` on PyPI, and the VS Code extension on the Marketplace. I held off on a broader announcement until the trust-system surface was coherent enough to talk about as one thing. The governance waveplan — column classification, per-env masking, 8-field audit trail on every run, `rocky compliance` rollup, role-graph reconciliation, retention policies — landed end-to-end last week in engine-v1.16.0 and rounded out in v1.17.4 (tagged 2026-04-26). That's the milestone I'd been waiting for. The pitch: keep Databricks or Snowflake. Bring Rocky for the DAG. Rocky is a Rust-based control plane for warehouse pipelines. Storage and compute stay with your warehouse. Rocky owns the graph — dependencies, compile-time types, drift, incremental logic, cost, lineage, governance. The things your current stack can't give you because it doesn't own the DAG. A few things I think are interesting: - Branches + replay. `rocky branch create stg` gives you a logical copy of a pipeline's tables (schema-prefix today; native Delta SHALLOW CLONE and Snowflake zero-copy are next). `rocky replay <run_id>` reconstructs which SQL ran against which inputs. Git-grade workflow on a warehouse. - Column-level lineage from the compiler, not a post-hoc graph crawl. The type checker traces columns through joins, CTEs, and windows. VS Code surfaces it inline via LSP. - Governance as a first-class surface. Column classification tags plus per-env masking policies, applied to the warehouse via Unity Catalog (Databricks) or masking policies (Snowflake). 8-field audit trail on every run. `rocky compliance` rollup that CI can gate on. Role-graph reconciliation via SCIM + per-catalog GRANT. Retention policies with a warehouse-side drift probe. - Cost attribution. Every run produces per-model cost (bytes, duration). `[budget]` blocks in `rocky.toml`; breaches fire a `budget_breach` hook event. - Compile-time portability + blast radius. Dialect-divergence lint across Databricks / Snowflake / BigQuery / DuckDB (12 constructs). `SELECT *` downstream-impact lint. - Schema-grounded AI. Generated SQL goes through the compiler — AI suggestions type-check before they can land. What Rocky isn't: - Not a warehouse — it's the control plane on top. - Not a Fivetran replacement. `rocky load` handles files (CSV/Parquet/JSONL); for SaaS sources use Fivetran, Airbyte, or warehouse-native CDC. - Not dbt Cloud — no hosted UI, no managed scheduler. First-class Dagster integration if you need orchestration. Adapters: Databricks (GA), Snowflake (Beta), BigQuery (Beta), DuckDB (local dev / playground). Apache 2.0. I'd love feedback on the trust-system framing, the governance surface (particularly classification-to-masking resolution in `rocky compile` and the `rocky compliance` CI gate), the branches/replay design, the cost-attribution primitives, or anything else that catches your eye. Happy to go deep in the thread. --- top comments --- [Xiaoher-C] The compile-time lineage part is the most interesting bit to me. A lot of “data lineage” tools feel like archaeology after the fact: parse logs, reconstruct what probably happened, then hope it matches reality. Having the compiler know “this column flows into these downstream models” before execution changes the workflow quite a bit. It makes refactors and masking policies much less scary. Do you expose any kind of “lineage diff” between branches? For example: this PR changes the downstream impact of `customer.email` from A/B/C to A/B/D. That would be useful in code review. [ramon156] If your introduction message already includes a bunch of uncurated claims and LLM smells, then what does that say about the code I'm about to run? [mollerhoj] Its a bit confusing to claim that "The things your current stack can't give you because it doesn't own the DAG" and use DataBricks as your example: DataBricks inclu […truncated]
Migrate SSRS reports from Snowflake to databricks..!
We are being onboarded on project (SSRS reports from Snowflake to databricks) have never worked on similar thing. As we were mostly in support role. So, can you guys please guide us how to approach this project. And what thing needed to be take care of. And if anyone worked on similar thing can you guide us with the rough process so that we can get a broad idea and move further. Thanks!
Databricks vs. Snowflake Weekly
--- top comments --- [noashavit] It’s hard keeping up with leading players in a fast-paced market like the data. I built this artifact to scrape relevant sources (docs, site pages, press release, etc) to surface the most recent, trusted and relevant news I should know about (with clear priority for product updates) Use it to keep up with these or other leading players in an industry you pay close attention to.
Snowflake vs. Databricks Showdown
Show HN: Altimate Code – Open-Source Agentic Data Engineering Harness
I'm Anand, co-founder and CTO of Altimate AI. My co-founder Pradnesh and I are open-sourcing Altimate Code. AMA. Why we built this: Pradnesh and I have been building tooling for data engineers for three years: dbt Power User and Datamates vscode extensions with combined 750k+ installs, running against real Fortune 500 data estates. The pattern we kept seeing: general-purpose agents can write SQL, but they have no model of what the SQL does. No lineage. No schema context. No understanding of what's in a dbt manifest. That's not a prompt problem; it's a missing tool layer problem. The numbers make it concrete: 27–33% of AI-generated SQL references tables that don't exist. 78% of errors are silent wrong joins, queries that compile, run, and return confidently incorrect data. One team got a $5k bill from a single Cortex AI query their resource monitors never caught. This isn't a model quality problem. It's a missing harness problem, and we proved it. Claude Code and Cursor are genuinely good for software engineering. But when you point them at a data stack, they hallucinate column names, ignore partition keys, and have no concept of data contracts or quality rules in your models. From building tooling against real data estates, we knew exactly what was missing at the tool level. We forked from OpenCode for the agentic scaffolding. What we added is the entire data layer: compiled Rust engines, purpose-built skills, and the harness that wires them together. What Altimate Code does that general agents can't: - Live column-level lineage: traces any column through joins, CTEs, and subqueries deterministically. 100% edge match on 500K benchmark queries at 0.26ms/query, and not from a cached manifest. Manifests go stale within hours on active pipelines, which makes cached lineage unreliable for anything agentic - SQL anti-pattern detection: 26 rules, zero false positives, 0.48ms/query. - Local SQL validation: interrogates your schema catalog in 2ms without touching your warehouse. Wrong table? Caught with a fuzzy-matched fix suggestion before the LLM goes into a fix loop. That's 10ms for 5 fix cycles vs. 2.5 minutes of Snowflake round-trips - Purpose-built skills for dbt development, testing, troubleshooting, documentation, SQL optimization, and migration. - 3 agent modes with compiled permission enforcement (Builder, Analyst, Planner). “Analyst” enforces read-only at the engine level, not just the prompt. That distinction is what makes it safe to run against production - Persistent memory: cross-session, two scopes (global preferences + project knowledge). Versioned in git, team-inherited on git pull - PII detection, SQL injection scanning, permission enforcement — all at the engine level, not the prompt - 10 data connectors: Snowflake, BigQuery, Databricks, PostgreSQL, Redshift, DuckDB, MySQL, SQL Server etc - Local tracer: every LLM call, tool invocation, and warehouse credit traced locally. No external services. On the benchmarks: We ran ADE-bench, the open standard from dbt Labs Altimate Code (Sonnet 4.6) → 74.4% Cortex Code (Snowflake) (Opus 4.6) → 65% Claude Code (baseline) (Sonnet 4.6) → ~40% A cheaper model with compiled tools outperformed a more expensive model without them. The gap is the harness. Full methodology is in the launch post and linked from the README. To try it: - npm install -g @altimateai/altimate-code - altimate - altimate /discover /discover interrogates your dbt projects, warehouse connections, and installed tools automatically. GitHub: https://github.com/AltimateAI/altimate-code · Docs: [altimate-code.sh](http://altimate-code.sh) There's a /feedback command that files a GitHub issue directly. If something breaks or doesn't behave the way you'd expect, use that or reply here. I'll be in this thread. --- top comments --- [mathisd] Few observations related to data engineering in the context of a data warehouse: 1. Protocols and IR (Intermediate Representation) have layed and contin […truncated]
Show HN: We built a billion row spreadsheet
We're former AWS S3 engineers. Row Zero is "Google Sheets on steroids". We built it because we were frustrated with Excel's 1 million row limit and poor integration with cloud data sources like S3, Postgres, Snowflake, and Databricks. Under the hood, the spreadsheet engine is a columnar key-value store written in Rust. We built it from scratch because SQL engines like DuckDB only work well for structured tables. The unstructured nature of spreadsheets make them a bad fit for SQL. The engine also needs to be 100% Excel compatible so you can import existing xlsx files. The front-end is a big React app with hand-rolled canvas rendering. We virtualize scrolling to handle 2 billion+ row data sets. All workbook state runs in the cloud (located in a data center close to you to keep things snappy). We launched on HN 2 years ago and got lots of great feedback. Since then, customers have pulled us towards the enterprise. Today we're launching Row Zero 2.0 with support for all the enterprise bells and whistles: OAuth database connections, SSO, SCIM, workspaces, private link on AWS and Azure, bring your own AI key, and private storage. Excited to hear what you think.
Show HN: Forge, the NoSQL to SQL Compiler
https://forge.foxtrotcommunications.net/ I've been a data engineer for years and one thing drove me crazy: every time we integrated a new API, someone had to manually write SQL to flatten the JSON into tables. LATERAL FLATTEN for Snowflake, UNNEST for BigQuery, EXPLODE for Databricks — same logic, different syntax, written from scratch every time. Forge takes an OpenAPI spec (or any JSON schema) and automatically: 1. Discovers all fields across all nesting levels 2. Generates dbt models that flatten nested JSON into a star schema 3. Compiles for BigQuery, Snowflake, Databricks, AND Redshift from the same metadata 4. Runs incrementally — new fields get added via schema evolution, no rebuilds The key insight is that JSON-to-table is a compilation problem, not a query problem. If you know the schema, you can generate all the SQL mechanically. Forge is essentially a compiler: schema in, warehouse- specific SQL out. How it works under the hood: - An introspection phase scans actual data rows and collects the union of ALL keys (not just one sample record), so sparse/optional fields are always discovered - Each array-of-objects becomes its own child table with a hierarchical index (idx) linking back to the parent — no manual join keys needed - Warehouse adapters translate universal metadata into dialect-specific SQL: BigQuery: UNNEST(JSON_EXTRACT_ARRAY(...)) Snowflake: LATERAL FLATTEN(input => PARSE_JSON(...)) Databricks: LATERAL VIEW EXPLODE(from_json(...)) Redshift: JSON_PARSE + manual extraction - dbt handles incremental loads with on_schema_change='append_new_columns' The full pipeline: Bellows (synthetic data generation from OpenAPI specs) → BigQuery staging → Forge (model generation + dbt run) → queryable tables + dbt docs. There's also Merlin (AI-powered field enrichment via Gemini) that auto-generates realistic data generators for each field. I built this because I watched teams spend weeks writing one-off FLATTEN queries that broke the moment an API added a field. Every Snowflake blog post shows you how to parse 3 fields from a known schema — none of them handle schema evolution, arbitrary nesting depth, or cross-warehouse portability. Try it: https://forge.foxtrotcommunications.net Happy to answer questions about the architecture, the cross-warehouse compilation approach, or the AI enrichment layer. --- top comments --- [Shyaamal11] The cross warehouse portability problem you're solving is real. I've watched teams maintain four separate FLATTEN implementations for the same pipeline just because they were multi-cloud. The compiler framing makes sense. Curious how you handle schema drift at the introspection phase specifically when an API starts returning a field as sometimes a string and sometimes an object depending on the endpoint response. Does Forge pick a winner or surface it as a conflict for the user to resolve?
Use RSA key snowflake connection options instead of Password
I want to connect to a Snowflake database from the Data Bricks notebook. I have an RSA key(.pem file) and I don't want to use a traditional method like username and password as it is not as secure as it exposes the password. pem_file_path = 'rsa_key.pem' with open(pem_file_path, 'r') as pem_file: pem_private_key = pem_file.read() # Configure Snowflake options sfOptions = { "sfURL": sfurl, "sfUser": sfuser, "sfDatabase": sfdatabase1, "sfSchema": sfschema, "sfWarehouse": sfwarehouse, "sfRole": sfrole, "pem_private_key": pem_private_key # Direct content of your unencrypted PEM file } # Establish your Snowflake connection (example, assuming you use Spark) from pyspark.sql import SparkSession spark = SparkSession.builder \ .appName("SnowflakeConnection") \ .getOrCreate() df = spark.read.format("snowflake") \ .options(**sfOptions) \ .option("dbtable", "your_table_name") \ .load() df.show()
NewsBayada’s Snowflake-to-Databricks Migration: Transforming Data for Speed & Efficiency
Bayada Home Healthcare migrated from multiple fragmented data platforms (SQL Server, Snowflake) to Databricks to unify data silos and enable self-service analytics across their 33,000-person organization. The 5-month migration, accelerated by Tredence's pre-built accelerators, positions Bayada to deploy AI agents for clinical documentation, workforce optimization, and automated compliance processes at their office level.
CommunityHow I Mastered Data Modeling Interviews
Data modeling structures business requirements into fact tables for transactions and dimension tables for context, enabling queries to answer business questions. Acing data modeling interviews requires mastering Kimball's dimensional modeling concepts like star/snowflake schemas and slowly changing dimensions, then systematically practicing diverse real-world examples by identifying business processes, events, entities, and table attributes.
Delta Lake 3.3.0
Delta Lake 3.3.0 adds Identity Columns for automatic unique keys, VACUUM LITE for faster transaction log-based cleanup, and enables Row Tracking backfill on existing tables for row-level lineage tracking. UniForm Iceberg can now be enabled on existing Delta tables without data rewriting, and Type Widening is now supported in Delta Kernel for reading type-evolved tables.
NewsHow Coinbase Built and Optimized SOON, a Streaming Ingestion Framework
Coinbase built SOON, a configuration-driven streaming ingestion framework using Spark Structured Streaming to unify table replication and Kafka ingestion into Delta Lake. The framework optimizes merge performance through techniques like min-max range filtering, k-means clustering, and change data feed synchronization for downstream systems like Snowflake.
Get Tuesday's version of this
Tracking Snowflake? The Tuesday email carries what moved across the whole ecosystem, not just this topic. Free, one-click unsubscribe.



