Databricks: From Spark to AI Agents—The Story Behind the $190 Billion Valuation

trade4wks agorelease AiFun
290 0

In August 2026, Databricks announced the completion of The $5 Billion Strategyfinancing, with a valuation of $190 billion.

This figure is already impressive enough, but what’s even more noteworthy than the valuation is its growth rate: Databricks’ annualized revenue run rate has already exceeded $7 billion, representing year-over-year growth of more than 80%, while maintaining positive adjusted free cash flow.

At first glance, Databricks could easily be seen as just another super-unicorn capitalizing on the AI boom.

But what makes it truly interesting is that, over the past decade or so, it has continued to expand along the same central theme:

From big data computing to data platforms, and on to enterprise AI infrastructure.

What it aims to do today goes beyond simply “helping companies analyze data”; it aims to become the future. AI Agent The data foundation required for operation.


I. Where Did Databricks Come From?

Databricks was founded at UC Berkeley's AMPLab.

Its earliest core technology was what later became the famous Apache Spark.

In the Hadoop era, big data processing relied heavily on MapReduce, and many computational tasks required frequent disk reads and writes, which was not very efficient.

Spark's breakthrough lies in:

In-memory computing significantly improves the efficiency of large-scale data processing.

In 2013, Databricks was founded by key members of the Spark team, including Ali Ghodsi, Matei Zaharia, Reynold Xin, and Ion Stoica.

This has also laid the foundation for the company’s long-standing product philosophy:

First, establish a technical presence through open-source projects, and then build a commercial platform centered on enterprise-level needs.

Spark is just the first step.


II. It Was Lakehouse That Truly Changed My Fate

Spark has been very successful, but it is open source.

Databricks must answer one question:

If Spark is free, why would customers pay Databricks?

In the end, the answer turned out to be Lakehouse.

In the past, enterprise data architectures were typically divided into two categories:

One type is Data Warehouse... It offers good performance and is well-established, but it is expensive and not suitable for handling large volumes of unstructured data.

Another category is Data Lake...It is low-cost and highly scalable, but its data quality, transactional capabilities, and governance capabilities are relatively weak.

The core concept behind Databricks’ Lakehouse is very simple:

Combine the low cost and openness of a data lake with the reliability and performance of a data warehouse.

This became the most important strategic foundation for the company over the next decade.


III. What Exactly Is a Technology Platform?

To understand Databricks, you can break it down into several key technical layers.

1. Delta Lake: Giving Data Lakes Database Capabilities

The problem with traditional data lakes is that, while data is stored cheaply, they lack transactions, version control, and consistency.

Delta Lake adds a transaction log on top of Parquet files, thereby providing capabilities such as ACID transactions, version rollback, and data consistency.

It addresses a fundamental issue:

Can companies build a reliable data platform directly on top of low-cost object storage?

Databricks' answer is: Yes.


2. Photon: Bridging the Gap in High-Performance SQL

Relying solely on Spark, it’s difficult to compete head-to-head with data warehouses like Snowflake and BigQuery over the long term.

So Databricks developed Photon in-house.

Photon is a high-performance vector query engine primarily used to accelerate SQL and analytical tasks.

It represents a significant shift in Databricks' business model:

The underlying data format can be open, but the high-performance computing engine can create a proprietary barrier.


3. Unity Catalog: Turning Data Governance into a Core Competency

In the age of AI, an increasingly important question is:

Who has access to what data?

Which model uses which data?

Does the AI Agent have permission to access customer information?

That is precisely where Unity Catalog's value lies.

It provides centralized management of data, models, permissions, auditing, and data lineage.

In the past, governance primarily served the people.

In the future, governance will likely serve AI agents primarily.

This is because an employee might access data dozens of times a day, while an agent might initiate tens of thousands of calls a day.

As a result, “permissions and governance” have gradually evolved from an IT management issue into an AI infrastructure issue.


IV. Transitioning from a Data Company to a Data + AI Company

Following the boom in generative AI, Databricks quickly realized:

What is truly scarce for companies is not the underlying model, but their own proprietary data.

GPT, Claude, Gemini, and Llama can all be replaced.

However, customer records, transaction data, product knowledge, supply chain information, and internal documents cannot be easily replicated.

Databricks is already positioned right on top of this data.

Therefore, in 2023, it reached approximately $1.3 Billion Acquisition of MosaicML...is officially expanding into large-model training and generative AI infrastructure.

The strategic logic is also very clear:

In the past, we managed:

Data → Analytics

Current coverage:

Data → Model → Inference → Agent


V. Why Did I Buy Another Neon?

Databricks has traditionally been strongest in analytical data systems.

It's very good at answering:

Which customers bought what last year?

However, once the AI agent is fully integrated into the business system, it may need to:

Create an order for this customer.

The former is analysis; the latter is trading.

So in 2025, Databricks acquired Neon, a serverless PostgreSQL company, and launched Lakebase.

This step is crucial.

This is because Databricks has begun to shift from “analyzing the past” to “directly participating in business operations.”

In other words, it not only aims to serve BI professionals and data scientists, but also hopes that the real-time databases for future AI agents will run on its own platform.


VI. What is Databricks, essentially, today?

If we were to simplify today’s Databricks into a single diagram, it could be roughly understood as:

Underlying Storage
S3 / Azure Data Lake / Google Cloud Storage

Open Data Formats
Delta Lake / Iceberg / Parquet

Computational Engine
Spark / Photon / Serverless

Data Engineering
ETL / Streaming / Lakeflow

Data Governance
Unity Catalog

Analytics and BI
Databricks SQL / Genie

AI and Models
MLflow / Mosaic AI / Model Serving

AI Agents and Business Systems
Agent Bricks / Lakebase / AI Gateway

So, Databricks is no longer just a “big data company.”

What it really wants to be is:

A unified operating platform for enterprise data and AI.


VII. Why is the market willing to value the company at $190 billion?

There are three main reasons.

First, Databricks specializes in the data assets that are most difficult for enterprises to migrate.

Models can be replaced, GPUs can be replaced, and even the cloud can be replaced.

However, the data, access permissions, data lineage, and business semantics that companies have accumulated over the years make migration extremely costly.

Whoever masters this aspect will enjoy strong customer loyalty.

Second, it is tapping into multiple markets at the same time.

In the past, Databricks primarily competed in the data warehouse market.

It has now entered:

  • Data Engineering
  • Data Warehouse
  • Data Governance
  • machine learning
  • Generative AI
  • PostgreSQL
  • AI Agent

As a result, its TAM continues to expand.

Third, AI agents may give rise to entirely new infrastructure requirements.

The workflow for traditional software is as follows:

Human → Application → Database

In the future, it might become:

Human → Agent → Tools → Database

In the future, a single employee may have multiple agents at the same time.

As the number of agents grows rapidly, data access, model calls, database transactions, and permission management will all increase accordingly.

Databricks is betting on this trend.


VIII. The competition between it and Snowflake has also changed

In the past, Databricks and Snowflake were considered direct competitors.

Snowflake is expanding from data warehousing into AI.

Databricks is expanding from data lakes and Spark to data warehouses and AI.

In the end, both sides headed in the same direction:

Unify enterprise data and AI workloads.

However, Databricks places greater emphasis on an open ecosystem.

The core logic is:

Data remains open, while computing, governance, and AI platforms are commercialized.

This is particularly appealing to large enterprises concerned about vendor lock-in.


IX. The Greatest Opportunity: Becoming the Data Operating System of the Agent Era

The real story behind Databricks' next phase may no longer be the Lakehouse.

Instead:

AI Agent Infrastructure.

In the future, a large enterprise may have tens of thousands of agents running simultaneously.

All of these agents require:

  • digital
  • Context
  • Memories
  • mould
  • Permissions
  • Governance
  • comprehensive database

Databricks is trying to bring all of these capabilities together on a single platform.

Lakebase handles the real-time database.

Genie is responsible for understanding enterprise data and business semantics.

Unity Catalog handles permissions and governance.

Mosaic AI is responsible for the model.

AI Gateway handles model invocation and management.

If this line of reasoning holds true, Databricks“ role will evolve from that of a ”data analytics platform” to:

An enterprise-grade runtime platform for AI agents.


X. Of course, the risks are also significant.

A valuation of $190 billion also indicates that the market has very high expectations for Databricks.

It faces at least three risks.

First, there are cloud giants like AWS, Microsoft, and Google.

They combine cloud, storage, databases, models, and GPUs, giving them a natural advantage in integration.

Second is Snowflake.

The product boundaries between the two companies are increasingly overlapping, and competition has expanded from data warehousing to AI.

Third, will AI platforms ultimately be absorbed by model vendors?

If, in the future, OpenAI, Anthropic, or cloud providers integrate agents, databases, and model governance into a single platform, some of Databricks’ value may be eroded.

Conversely, if enterprises adhere to a multi-cloud, multi-model, and open data architecture, Databricks’ position as a neutral platform will become even more important.


Conclusion: Databricks’ true moat is data gravity.

Looking back at Databricks’ growth trajectory, it’s clear that every step has centered on the continuous expansion of data.

Spark handles the computations.

Delta Lake provides reliable storage.

Lakehouse bridges the gap between data lakes and data warehouses.

Photon resolves performance issues.

Unity Catalog addresses governance.

Mosaic AI Model Solutions.

Lakebase handles real-time transactions.

Agent Bricks and AI Gateway, meanwhile, have begun addressing the challenges of deploying AI agents in production environments.

All of these capabilities ultimately revolve around a single core:

Data Gravity.

Computing will move closer to where the data is.

Wherever computing goes, AI will follow.

So, behind the $190 billion valuation, what investors are really betting on may not be a stronger data warehouse company.

Instead:

Can Databricks become the unified operating system for enterprise data, models, and agents in the AI era?

© Copyright notes

Related posts

No comments

none
No comments...