Reference · Chapter 38
Databricks Versus Snowflake
The question arrives as an architecture question and it is almost never one. Somebody's Snowflake bill has grown faster than the business did, a vendor has explained that a lakehouse would be cheaper, and the meeting is now about platforms when it started about money.
That is worth separating before anything else, because the honest answer depends entirely on which question is being asked.
The short version
Snowflake and Databricks are both excellent, and for ordinary SQL analytics they will feel similar to the people using them. The difference shows up in what else you need. Snowflake is a warehouse that has grown outward into streaming and machine learning; Databricks is a processing platform that has grown inward toward warehouse-grade SQL. If your work is dashboards and queries, the warehouse fits. If your work has become long-running transformation, model training, and applications built on the data, the platform fits. What decides most migrations, though, is a bill, and a bill is driven by the shape of the workload rather than by the logo on the invoice. Move the same workload unchanged and you mostly move the invoice.
What actually drives the bill #
Most teams assume the expensive part is the heavy aggregation. Usually it is not.
One consultancy that does this work for a living put the pattern plainly: "the primary driver of Snowflake costs is not the compute for aggregation, but the compute required for lots of reads/scans." Their fix for one customer was to keep hourly aggregates in Snowflake and push the results into a store built for serving, so the dashboards stopped hitting the warehouse. That "slashed their Snowflake bill by ~25%, which was worth millions to them."
Read that carefully, because it did not involve leaving Snowflake. It involved noticing that user-facing dashboard traffic and analytical aggregation are different workloads, and that paying warehouse prices to serve a dashboard is the actual defect.
Snowflake is good, and it is not a relational database #
The fairest assessment of Snowflake we have read came from an engineer who spent months trying to reduce a bill and came away more impressed than when he started. He describes gaining "respect as in 'this is really great' as well as respect as in 'I need to be on guard here or I'm going to get hurt.'"
His diagnosis of why teams overspend is the useful part. "I think my biggest misconception at the outset was thinking of Snowflake like it's a relational database. It's not." There are no b-tree indexes. There are clustering keys, which colocate data so queries can prune micropartitions, and a well clustered table filtered on its clustering key is fast. Join on something else and you will pay for it, unless you enable search optimization, which costs more again.
That is not a flaw so much as a different machine wearing familiar clothes. It does mean that a team applying Postgres instincts to Snowflake will produce a bill that looks like a platform problem and is actually a modelling problem. Migrating that team to Databricks moves the same instincts onto a platform with its own set of ways to be expensive.
Where the two genuinely differ #
Set cost aside and the real difference is the shape of the work.
Snowflake's compute model is built for spiky analytical SQL and it is very good at it. You pay per second with a sixty-second minimum when a warehouse wakes, auto-suspend is on by default, and a warehouse that is not running costs nothing. For a business that queries hard in office hours and barely at night, that model is close to ideal.
It fits less comfortably when the work stops being queries. Long-running transformation, Spark jobs and training runs hold compute for hours, and the thing that made short queries cheap does nothing for you there. That is the workload where Databricks was built first and where the argument for it stands on its own, without needing a cost story.
The second real difference is what happens past the dashboard. When the interesting work becomes software, an operator console, an agent with permissions, a service that writes back, a lakehouse with an operational database beside it does something a warehouse was never meant to do.
The migration that relocates the bill #
What to try before you move #
Three things, in order, and none of them needs a migration.
Find out where the money actually goes. Snowflake's own query history will tell you, and the answer is usually more concentrated than anybody expects. A small number of dashboards, or one badly clustered join, frequently accounts for a startling share.
Move the read-heavy serving off the warehouse. If user-facing dashboards are scanning all day, they belong somewhere built for that, and the aggregation can stay exactly where it is.
Then look at what remains. If what remains is spiky analytical SQL, you have a healthy warehouse workload and no reason to move. If what remains is hours of transformation, training runs and applications waiting to be built, you have found a real platform argument, and it is a better one than the bill was.
Worth knowing that this is easier than it used to be. Snowflake and Databricks now read and write each other's Iceberg tables, so moving a workload no longer means moving the data first. The same consultancy above noted the constraint in passing: running aggregations in Spark to cut costs works, "but you would then need Iceberg to make the tables queryable in Snowflake." That is exactly the thing both vendors shipped this year.
How to decide #
Name the workloads and cost each one where it runs today. Separate the SQL that serves dashboards from the transformation that runs for hours, from the training, from the applications somebody wants to build next year.
Then ask which platform each would sit on if you were starting fresh, and whether the split is worth two bills. That exercise decides itself more often than it looks like it will, and it decides against migrating more often than a migration vendor's website will tell you.
When it does decide the other way, the reasons are specific, and specific reasons are what keeps a project alive in month four when it gets hard.