As AI improves, data becomes more valuable and one company may be uniquely situated to take advantage of this.

Databricks was founded in 2013 by a team of seven researchers with the goal of “making big data processing more accessible.” Today, it has evolved into the fifth most valuable private company in the world.

AI is central to the Databricks story and the proof is in the valuation. Less than two years ago in December of 2024 Databricks raised $10 billion in a Series J funding round that valued them at $62 billion. In August, the company closed a massive strategic funding round valuing the company at $190 billion.

Databricks is one of the most important and largest private companies and yet at the same time due to its complexity it’s one of the least understood. By way of a roadmap, this piece begins with the history of the company before explaining what exactly the company does in terms that are understandable without a degree in data science. I will then speak about the leadership, ethical questions involved and cover some of the financing before finally looking at the TAM, competition (including their rivalry with Snowflake) and valuation. We will conclude the piece by going over final thoughts and discussing a potential IPO in the future.


Welcome to The Private Ledger - an independent research publication that focuses on breaking down private companies and macro trends surrounding IPOs.

  • 📈 You’re reading my latest breakdown into a private company. If you enjoy the research, please feel free to like & subscribe to receive much more like this every week!


The Birth of Databricks

Before we discuss Databricks, we have to take a step back to understand what problem Databricks came to solve.

The internet was revolutionary in numerous different ways but one of the most underrated parts about how the internet changed life as we know it is the amount of data that companies now have access to. Originally, this data was stored on a single computer that could process this data.

At some point around 2007, this system was no longer sufficient due to the amount of information that companies now had access to. Imagine Apple processing their entire database on one server — it simply wasn’t possible.

So, companies evolved and began to incorporate multiple machines that each divided the data and then worked together in order to process data simultaneously. This data would subsequently be collected and compiled to process enormous amounts of data at the same time, something a single server wasn’t capable of doing. As more data came in, companies simply had to add more servers to deal with the additional data. This system was done through an open source software called Hadoop.

However there were a number of problems with Hadoop including relying on physical disk drives which are slower, more fragile, and significantly less capable for machine learning workloads. So a group of researchers from UC Berkeley decided to make a fundamental change.

The researchers decided to move the data to the memory of the device itself instead of physical disks. This system was known as Apache Spark and it became the system that Databricks still uses today. We’re going to come back to this shortly but I don’t want to get too technical yet.

And so, the creators of Apache Spark realized that many people wanted this whole process done for them. And so, they created Databricks.

What does Databricks do?

Despite the size of Databricks, every explanation I found seemed to be written by data engineers for data engineers. So I’m going to start by simplifying the company as much as possible and expanding out from there.

Simply put, Databricks is a massive data storehouse that allows companies to store, clean, manage and analyze huge amounts of data. In the past few years, companies have now begun using AI and machine learning to speed up the process of analyzing and receiving insights from data.

Think of it as the warehouse that sits in between a company’s huge amounts of data and AI and machine learning algorithms trying to understand and utilize it.

Now, let’s return to Apache Spark. Prior to Apache Spark, companies had to store their data in two separate storage systems. The first was a traditional data warehouse that stored mostly structured data such as spreadsheets and other relatively simple conventional types of data. The second was a “data lake” that stored more complex unstructured data like videos or an email thread.

To fix this, the creators of Apache Spark launched Databricks, which combined both of these and pioneered a new idea called a “Lakehouse,” allowing companies to store the data in one unified place.

Interestingly, when Databricks was founded, they decided to keep Apache Spark open source - making it free for companies to use. And use they did. The world spent the next dozen years building technology on top of Apache Spark. Today, Apache Spark’s website boasts that thousands of companies including 80% of Fortune-500 companies use Apache Spark.1

Data Lakehouse Architecture | Databricks

Matei Zaharia, the initial creator of Apache Spark spoke about this decision in a recent interview:

“One of the reasons to open source something is if you think it’s a layer that there will be some network effect and it will benefit from many people collaborating on it.”

Databricks essentially gave their software away for free to the world. So what was their product?

Databricks offered companies proprietary features and an option to manage and maximize its data, the trust stemming from the fact that these were the same researchers who originally built Apache Spark.

Many companies that didn’t want to manage highly complex data decided to opt for this approach and began to use Databricks’s service to manage their data. And once they began to use Databricks, they usually became long-term loyal customers.

While there hasn’t been a churn rate released publicly, in 2021, Databricks announced that their net retention rate was 150% meaning customers spent 50% more than they had spent the previous year even factoring for churn.2 Meaning not only did customers likely churn at a very low rate, they actually increased their usage and spending. In 2024 and in 2025, that number stayed consistently at 140% even with the company growing to billions in revenue.34

The AI revolution:

After a few years of successfully running Databricks, the “AI revolution” happened. However, rather than many companies that were threatened by AI, Databricks was actually in a prime position to capitalize on it.

“Databricks has spent a decade being early to where AI was headed. Now it’s the infrastructure the industry builds and scales AI on,” - Thomas Laffont5

Much of this happened because of Matei Zaharia and his decision to open source Apache Spark. Because huge amounts of data were being stored on the Apache Spark software, Databricks essentially had much of the work outsourced for them, creating a moat and positioning themselves as the leaders in data management.

“Imagine our thing wasn’t open and there is an open one, which one is going to win in the long run?…We benefit because we don’t have the time to write connectors to a thousand different data bases and file formats but we can just use the ones people make and of course they benefit from joining.” - Matei Zaharia

And so when AI reached mass adoption, Databricks was in prime position to capitalize.

AI was able to train and run on Databricks lakehouses without having to move the data to a separate location. This allowed Databricks to make the product significantly more efficient while training and improving AI models at the same time. And so the existing infrastructure that Databricks had owned and operated for the previous decade just became even more valuable.

By allowing companies to use their product, the researchers who built Apache Spark and founded Databricks became the de facto industry-leading experts for companies that were looking for someone to run the data analysis for them. With AI as a massive tailwind, Databricks continued to improve their product, revenue expanded and the valuation tripled.

The Leadership:

Ben Horowitz, the co-founder of a16z, one of the most successful venture capital firms in the world, called Ali Ghodsi the best CEO in Silicon Valley today. And yet, Ali Ghodsi never wanted to be the CEO.6

Back in 2015, Databricks was in an interesting place. Apache Spark was extremely successful and yet, Databricks was struggling with commercial success.7 And so, Ghodsi actually decided to follow his dream and apply to become a professor at UC Berkeley. Instead, Databricks asked Ghodsi to become the CEO. Ghodsi said the decision was a hard one, but what swayed him eventually was that choosing the more challenging option in his life had previously opened doors and he figured the same would be true with the CEO position. Today, it’s become a philosophy of his:

“I want to challenge myself as much as possible, after all, how many years do we have on this planet?”
Ali Ghodsi - Databricks CEO.

Unbeknownst to Ghodsi, this CEO appointment was only a six-month trial run. Over a decade later, the trial has clearly proved to be a success. In the decade under the leadership of Ghodsi, Databricks has gone from a tiny company generating only $1 million in revenue to an enormous, rapidly growing behemoth that generated billions of dollars in 2025.

Ghodsi isn’t the only one responsible for Databricks’ success in the past year. There are a number of other vital members of the Databricks team including but not limited to: Matei Zaharia, Reynold Xin, Arsalan Tavakoli-Shiraji, Ion Stoica, Andy Konwinski, Patrick Wendell and David Conte,.

Databricks Success Story | Founders | Business Model
The Databricks founding team, from left to right Arsalan Tavakoli-Shiraji, Ion Stoica, Andy Konwinski, Reynold Xin, Ali Ghodsi (front center), Patrick Wendell, Matei Zaharia (front right).

Matei Zaharia is the aforementioned creator of Apache Spark. While working as a PhD student at only 24 years old, Zaharia created the program in order to solve the speed limitations and processing inefficiencies of the Hadoop system. A few years later, he joined the initial seven as one of the founders of Databricks. Today, he continues to lead Databricks’ technical side forward and is their CTO.

Reynold Xin originally built a SQL (Structured Query Language) layer on top of Apache Spark that essentially compiled physical queries into Spark programs and allowed Apache Spark to process data 100 times faster. Today he is still with the company under the title Chief Architect building their cloud computing system.

A final crucial member of the team that is worth noting is David Conte, who joined Databricks in October 2019 as the CFO and has since become an integral part of the Databricks team. Today, he is credited with much of Databricks’ success including leading all operational and financial functions through the company’s rapid growth over the past seven years.

Ethical Questions:

As with any company that I write about, it’s important to look at the ethics behind the company rather than simply the dollar amounts. With Databricks, the ethical questions are relatively minor.

Think of Databricks as an empty storage warehouse with no external users. The actions of the user inside the warehouse, as long as they are not damaging people outside of the warehouse, are entirely permissible and contained. This means that as long as Databricks users are following their terms and conditions (which limit activities like cyberhacking and human rights violations), what is done with the data is entirely up to the user and Databricks is not responsible for it.

Importantly, Databricks doesn’t have access to the data of the users renting their workspace due to the fact that they operate under a strict “zero trust” architecture.

The logical question follows, in that case, how does Databricks know when people are breaking their terms of service?

The answer is that while Databricks can’t monitor what is done inside of their warehouses, they can control what exits them. If they flag a user who uses a Databricks account to try and hack other sites, they can instantly shut down their account. The same goes for mining cryptocurrency or other actions that are against their terms of service.

Next, a Databricks subsidiary has been accused of using pirated books to train their AI models. While Databricks originally beat the claims in court, further lawsuits in late 2025 and 2026 have proceeded and are going to be heard by courts later this year.89

A final potential ethical question that arises is the choice to keep the company private for so long. Databricks is now worth almost $200 billion and generates billions in revenue every year and yet retail investors have zero access to it. This sets a dangerous precedent that is becoming increasingly popular right now. Institutional investors get the gains, retail gets left with the leftovers.

While as a retail investor I understand the frustration of feeling like I am being left with corporate leftovers, I don’t think this constitutes a legitimate ethical question. Ali Ghodsi has repeatedly said that they want to IPO, but not under the current market conditions.

Ultimately, while no company is perfect, the major ethical questions that surround Databricks are more about the usage of their customers rather than the platform itself. I believe that with Databricks you can invest with a clear ethical conscience.

TAM:

In August of 2025, Ali Ghodsi said that the database TAM is $105 billion and has been sitting there, untouched for the past 40 years. He continued and said that the way for Databricks to crack into that TAM was actually by continuing to embrace AI.10 Note that this figure is only referring to Databricks operational database market which is a newer market following their Neon acquisition.

During that same speech in 2025, Ghodsi said that from what he observed, a year ago, (i.e., 2024) 30% of databases were created by AI rather than humans. This year (i.e., 2025) 80% of databases were created by AI. He expected that by 2026 that number would grow to 99%. And that, he said, was the key to unlocking Databricks’ potential TAM.

“There’s a new user. The user is not human. It’s an AI agent, and if we just double down on making that user persona successful, that’s the wedge to disrupt that TAM.”- Ali Ghodsi, 2025.

However, even this massive TAM may be slightly outdated. Over the past year AI agents have used and created incredible amounts of data for complex tasks. More applications have been created, activity per app, new data types and more, all need to be stored somewhere. In other words, AI doesn’t simply help Databricks achieve their TAM; rather, AI actually expands Databricks’s potential TAM.

Unsurprisingly, the company best positioned to capitalize on that is Databricks. This isn’t by chance. While data can be traditionally stored in a simple database, AI specifically needs this data stored in a more complex fashion which Databricks has built its company around.

Lucky for Databricks, its TAM will continue growing irrespective of what they do. However, in order for Databricks to be able to capitalize on this expanding TAM, they need to be capable of working and innovating faster than the competition. Which naturally leads us to their competition and specifically one competitor, Snowflake.

The clearest rival is Snowflake. But the larger questions still remain. Is Databricks actually worth $190 billion and if so, is there still room for growth?

Below the paywall, I examine the relationship between Snowflake and Databricks and what Databricks needs to do for its valuation to make sense. I close with my view about a potential IPO including why it hasn't happened yet, when it might happen and whether Databricks would be attractive for retail investors.

Subscribe for the full breakdown.

Continue this report on Substack.

The preview above is available here. The complete analysis and references are for paid members on Substack.

Unlock on Substack ↗
Originally published in The Private Ledger. View original post.