Data insights

I benchmarked Databricks against my iPhone

October 7, 2026
I benchmarked Databricks against my iPhone
The device in your pocket easily runs most practical analytics workloads.

Big data is dead. Those of us who have spent the last decade working on data management systems know that real-world business data sets are much smaller than the ones that get talked about in benchmarks. What is less well understood is just how much faster everyday hardware has gotten. Workloads that once required a distributed system can now run on an iPhone. To prove that this isn't hyperbole, I went ahead and did just that.

My iPhone 17 Pro. Holds a lot of pictures of my dog, tackles giant datasets with ease.

The workload I chose was TPC-H. TPC-H is a set of 22 analytical queries against the database of an imaginary wholesale supplier.

TPC-H can be generated at different scales. I wanted to choose a scale that is representative of the high end of realistic business-user workloads. Snowflake and Amazon have published statistics about the real-world distribution of query sizes, which we can use to calibrate our benchmark. I ran at 4 scales, 25, 50, 100, and 200 GB, to approximate the high end of the real-world distribution.

I ran the queries on my phone using DuckDB, and on 3 different sizes of Databricks cluster, XS, S, and M, using Databricks Serverless SQL. DuckDB on my iPhone was faster than the Databricks clusters at all but the largest scale, and even there it was competitive.

This is surprising: even the smallest Databricks cluster has more CPU cores than an iPhone:
‍

Databricks XS Databricks S Databricks M iPhone 17 Pro
CPUs 24 40 80 6
Memory (GB) 128 224 448 12
Cost / Hour ($) 4.20 8.40 16.80 N/A

There are 2 reasons for these results. First, even though this workload is large by the standards of real-world business analytics, it's quite small compared to what modern CPUs can do. So the fixed costs of query planning and compilation are large in this benchmark. DuckDB's design excels at running small queries fast.

Second, DuckDB is a single-node database, while even a Databricks XS is a multi-node system. On every large JOIN and GROUP BY, Databricks has to perform a shuffle to distribute the data across nodes:

If your query can fit on a single node, it is more efficient to avoid these costs. There are very few queries that don't fit on a single node anymore: the AWS c9 series is available with 192 Graviton5 cores.

The most important implication of this finding is about cost. If you are using a system like Databricks or Snowflake simply to run SQL queries or Python dataframes against business data, you are paying a very high markup on the underlying compute.

Over the last 10 years, the cost of cloud compute has plummeted 10x, and the margins of the major data infrastructure providers have grown, resulting in the huge markups we see today. 

Running TPC-H on an iPhone is a fun stunt. Most companies aren't going to adopt an iPhone as their data warehouse — though if you do, I recommend you put it in an ice pack, or you'll lose about 30% of your performance to thermal throttling.

What this stunt shows us is that realistic business workloads are not at all challenging for modern computers. Most real-world workloads could run on single machines using execution engines like DuckDB and Polars that are designed to take advantage of the efficiencies of non-distributed execution. Importantly, using a single-node execution engine doesn't mean your entire company's workload has to run on a single machine. Queries from many users can be distributed across many workers.

This is how our Lake Compute service works: we have a large pool of worker nodes, but each worker node only works on one customer dbt model at a time. It's easy to assess whether your workload is a good fit for single-node execution engines: all the major data platforms now have built-in conversational analytics, so just ask!

‍"Show me a histogram of the size of data read by queries in my production data warehouse, with the x-axis on a log scale, units of GB."

I promise you will be shocked how small the vast majority of your queries are. We are all using expensive distributed execution engines for queries that could run on an iPhone. The way we are going to take advantage of cheaper compute is by moving to a new architecture where all the participants in the lakehouse talk directly to the storage layer.

You adopt this architecture in a stepwise, layer-by-layer process.

  1. Change your ingest to write to Iceberg tables. If you're a Fivetran user, you can use our Managed Data Lake migration workflow to convert your existing tables, including historical data, to Iceberg. The existing tables in your data warehouse will be transparently migrated into external tables so your queries continue to work. 
  2. Reconfigure your transformations to output to Iceberg. All the major compute engines support outputting to Iceberg, so this is mainly a matter of changing a setting. If you're a dbt user, you can use our Lake Compute service to execute dbt models. It uses DuckDB and single-node execution in its implementation.
  3. Move selected read workloads, like notebooks and ad-hoc queries, to use local compute. The cheapest CPU is the one on your desk!

Leveraging cheap and free compute isn't just about saving money. It means you no longer need to ration compute to your users. AI agents have given everyone their own analyst who can answer any question they can think of, but only if they aren't bottlenecked by sharing a small, expensive compute cluster. If we're going to connect AI to data, we're going to need open data infrastructure that leverages the cheap compute that's all around us, even in our pockets.

Details to reproduce this benchmark are in github.com/fivetran/iphone_benchmark

[CTA_MODULE]

Read more insights straight from the Fivetran + dbt CEO
Read more

Related blog posts

Start for free

Join the thousands of companies using Fivetran to centralize and transform their data.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.