Gunjo · Business Intelligence for the AI Era
← Sticker Wall JOURNEY · DETAIL

Databricks: From a 'Worst-Ever' Pitch Deck to a $188 Billion Data and AI Platform by Seven Berkeley Scholars

Founded: Ali Ghodsi (CEO since 2016), Ion Stoica (First CEO, Co-Director of Berkeley AMPLab), Matei Zaharia (CTO, Creator of Spark), Andy Konwinski, Patrick Wendell, Reynold Xin, Arsalan Tavakoli-Shiraji · Databricks, Inc.

JOURNEY

Key Fields

FIELD STAMPS
IndustryCloud Computing
RegionUS
ScaleGiant
ChannelOther

Origin

The story began with a $1 million competition. In 2006, Netflix offered a $1 million prize for improving recommendation algorithm accuracy by 10%. Berkeley PhD student Lester Mackey formed a team but struggled with massive data tools. Matei Zaharia, from the same lab, wrote a 600-line Spark engine in 2009 to help him iterate. Although the team matched the accuracy, they missed the prize by 20 minutes. While helping Facebook and Yahoo tune Hadoop, Zaharia realized MapReduce was disastrous for iterative machine learning calculations. Convinced the MapReduce architecture was flawed, he gathered six lab colleagues to refine Spark, separating storage from compute to accelerate iteration. They open-sourced it in 2010, donated it to the Apache Foundation in 2013, and the seven founders moved to San Francisco to commercialize the engine they built.

Milestones

2013
Worst Pitch Deck Growth
Most of the seven founders were Berkeley professors and PhD students with no business experience. When they pitched Ben Horowitz at a16z, the deck was described by Horowitz as having terrible graphics and ideas that hovered between condescending and insane—one of the least professional decks he had ever seen. Ghodsi later joked that the first pitch was a disaster. However, AMPLab colleague Scott Shenker had previously praised Zaharia to Horowitz as the most brilliant distributed systems scholar of the decade. Relying on his judgment of the people rather than the deck, Horowitz led the $13.9 million Series A, launching Databricks.
2016
Open Source Backlash Failure
The seven scholars assumed that open-sourcing Spark and building a community would lead to paid adoption. In reality, enterprises only wanted the free community version and refused to upgrade. Even Cloudera and Hortonworks reportedly refused to license Spark, forcing Databricks to commercialize it themselves. There was an earlier 'sliding door' moment: the Spark team wanted to license to Hortonworks but were rejected, forcing Databricks to pivot from an open-source community operator to a SaaS company selling cloud-hosted Spark. Commercialization was slow for the first three years; the academic team lacked sales and pricing expertise, relying on community buzz while struggling to generate revenue.
2016
CEO Transition Turning Point
The first CEO, Ion Stoica, a senior Berkeley professor and co-director of AMPLab, recognized that commercialization was not his strength and stepped down in favor of Ali Ghodsi, the former VP of Engineering and Product. Ghodsi, who grew up coding on a Commodore 64 and fled the Iranian Revolution with his family at age five, was the most decisive in product strategy and capital engagement. After the leadership change, Databricks consolidated its strategy into cloud-hosted Spark, abandoned the fantasy of being an open-source non-profit, and began charging based on compute resource usage rather than subscription fees.
2017
Azure Integration Turning Point
In 2017, Microsoft made Databricks a first-party service on Azure called Azure Databricks. It bore the Databricks name but was operated by Databricks—a model almost unheard of at the time for partnerships between tech giants and open-source startups. Ben Horowitz emailed Satya Nadella in 2016 to recommend it. Nadella grasped the critical importance of keeping data and processing software in the same place. Microsoft engineers adapted Spark for Azure and trained their sales force to sell it. Leveraging the Azure sales machine, Databricks' customer count scaled rapidly, marking the true inflection point for growth.
2019
Microsoft Investment Growth
Microsoft participated in Databricks' $250 million Series E at a $2.75 billion valuation. Horowitz noted that while Microsoft was feared and despised by startups in the 90s, Nadella completely redefined the relationship, turning Microsoft from a threat into a premier partner. This round established the unique landscape where AWS, Google Cloud, Microsoft, and Salesforce all hold stakes in Databricks, forming the capital foundation for its cross-cloud neutrality.
2020
Attacking the Warehouse Turning Point
In April 2019, they open-sourced Delta Lake, bringing ACID transactions and schema constraints to the data lake. In November 2020, they launched Databricks SQL (formerly SQL Analytics), officially entering the data warehouse market and introducing the 'Lakehouse' concept, which combined the flexibility of data lakes with the governance of data warehouses. This was the turning point where Databricks shifted from being indirectly adjacent to Snowflake to competing directly in the same market: ML engineers used Databricks for training, while BI analysts used Snowflake for queries—now, both were fighting for the same enterprise data budget.
2021
Benchmark War Turning Point
In 2021, Databricks released a benchmark report claiming the Lakehouse was 2.7x faster and 12x more cost-effective than Snowflake, implying that warehouse solutions become prohibitively expensive as data volume grows. Snowflake fought back with a two-week daily benchmark war, quickly adopted Apache Iceberg to replace Delta Lake, and launched Snowpark to pull data scientists into their ecosystem. The two companies were no longer just parallel products but were actively poaching customers and mindshare. By 2025, Databricks SQL reached over $1 billion in annual gross revenue.
2023
MosaicML Acquisition PMF
Databricks acquired MosaicML, a two-year-old generative AI startup with 60 employees, for $1.3 billion—a price tag of $21 million per employee, which the market deemed astronomical. However, this was a critical step in pulling generative AI into the data platform. Previously, Databricks only managed the data lake for training; the acquisition completed the chain for fine-tuning, training, and deploying enterprise-owned large models, allowing companies to train their own LLMs on their own data without exposing sensitive information. That same month, they integrated the flagship AI product into the platform, shifting Databricks' value proposition from saving ML compute costs to becoming enterprise AI infrastructure.
2024
Mega Funding Growth
Thrive Capital led a $10 billion Series J, pushing the valuation to $62 billion—one of the largest single VC investments in recent years. The NYT called it one of the most significant venture deals in history. The valuation narrative was rewritten from Big Data SaaS to AI infrastructure; investors preferred to bet on Databricks as the data foundation for the AI era rather than a pure generative AI play, valuing a data company at the level of cloud giants.
2025
Agents Infrastructure PMF
At the Data + AI Summit, they launched Agent Bricks (AI Agent development workbench), Lakebase (OLTP database for AI Agents), and Databricks One (no-code BI). By completing the operational database layer required for AI Agent execution, they repositioned the data platform as AI Agent infrastructure. That same month, they signed a $100 million, five-year partnership with Anthropic to integrate Claude, and in September, a $100 million deal with OpenAI to integrate GPT, allowing customers to call external models within Databricks' security perimeter without data leaving the domain.
2025
Series L $134B Growth
Insight Partners, Fidelity, and J.P. Morgan led a $4 billion Series L at a $134 billion valuation—the third major funding round in 12 months. During this period, they disclosed that ARR surpassed $4.8 billion (over 55% YoY growth), AI-specific revenue exceeded a $1 billion run rate, and free cash flow turned positive for the first time. An open-source-born data company officially entered the revenue tier of cloud giants.
2026
$188B Valuation Turning Point
In June, at the Data + AI Summit, they launched the open-source Omnigent Agent harness control plane, Lakehouse RT real-time engine, LTAP single-price transaction analysis architecture, Genie One agent colleague, and Genie ZeroOps self-healing operations, while announcing the acquisition of AI security operations platform Panther Labs. ARR grew 80% YoY to $6.9 billion. In July, Coatue led a $3 billion round, pushing the valuation to $188 billion. It is now the highest-valued Pre-IPO software company, and the market views Databricks as the primary anchor for the largest software IPO window in the next two years.

Turning Points

  • In 2016, Ion Stoica stepped down as CEO for Ali Ghodsi, shifting the strategy from operating an open-source community to selling cloud-hosted Spark—the first watershed moment for the seven scholars to transition from an open-source stance to commercial muscle.
  • In 2017, Microsoft made Databricks a first-party Azure service, leveraging Microsoft's global sales force to scale customers from thousands to tens of thousands. This deep integration, rare for open-source startups, laid the capital and channel foundation for cross-cloud neutrality.
  • In 2020, the launch of Delta Lake and Databricks SQL introduced the 'Lakehouse' concept, shifting the battlefield from the ML engineer niche directly into Snowflake's data warehouse home turf, turning indirect adjacency into direct budget competition.
  • In 2023, the $1.3 billion acquisition of MosaicML integrated generative AI into the data platform, upgrading the value proposition from saving ML compute to enterprise-owned LLM infrastructure, opening a window for re-pricing in the AI era.
  • In late 2024, the $10 billion Series J at a $62 billion valuation rewrote the narrative from Big Data SaaS to AI infrastructure. Capital preferred to bet on the data foundation rather than pure generative AI, moving Databricks into the cloud giant revenue tier.

Failures & Pitfalls

  • Open-source commercialization backlash: For the first three years, they assumed that open-sourcing Spark would lead to revenue. In reality, enterprises only wanted the free version. Commercialization was slow, the academic team lacked sales experience, and they struggled with cash flow, forcing a pivot from an open-source operator to a cloud-hosted SaaS company.
  • Worst pitch deck: The seven Berkeley scholars had never run a company. Their pitch deck to Ben Horowitz was described as having terrible graphics and ideas that were 'between condescending and insane.' Ghodsi himself admitted the first pitch was a disaster.
  • Sliding door moment: The Spark team wanted to license to Hortonworks but were rejected. If Hortonworks had accepted, Databricks might have remained just an open-source community operator rather than a cloud-hosted SaaS giant.
  • Pre-2017 data scientist trap: Before and around the launch of Delta Lake, Databricks was trapped in the niche of data scientists and ML engineers. The bulk of BI analyst budgets flowed to warehouse vendors, and for years, the company failed to capture the largest share of enterprise data budgets.
  • Model performance: After the MosaicML acquisition, while generative AI capabilities were complete, their proprietary models (Dolly and DBRX) could not outperform Llama 2 or Mixtral on public benchmarks. The proprietary model line struggled to compete as a cutting-edge AI story, leading Databricks to monetize through its platform rather than model performance.

关键成功要素

  • The seven founders were the original authors of Spark. The open-source community had already become the de facto big data standard before commercialization. Databricks was the natural commercial exit for the ecosystem, using open source to block competitors before monetizing.
  • Key components like Delta Lake, MLflow, and Unity Catalog were kept open-source, commoditizing expensive items like data catalogs, ML lifecycles, and governance. This focused the value on Databricks' primary revenue stream—compute usage—a classic platform strategy of lowering the price of complements.
  • Data stays in the customer's cloud, compute is charged on-demand, and they leverage AWS, Azure, and GCP sales channels. This avoids vendor lock-in from traditional warehouse vendors and mitigates market exclusion from any single cloud provider. Cross-cloud neutrality is both a technical architecture and a governance/capital structure.
  • The 2023 acquisition of MosaicML and the 2025 $100 million partnerships with Anthropic and OpenAI integrated external LLMs into their security perimeter. This allows enterprises to run Agents on private data without sending it to OpenAI, creating the stickiest data moat in the AI era.
  • Leveraging Microsoft Azure as a first-party service and Microsoft's 2019 investment opened enterprise channels. With AWS, Google Cloud, Microsoft, and Salesforce as shareholders and channels, Databricks is pushed to enterprise customers by multiple cloud sales forces—a rare structure for an open-source startup.

Lessons

  • Open source does not guarantee direct monetization; it often requires feeding the community before pivoting to hosted SaaS. Databricks' early struggles proved that running an open-source project and generating commercial revenue are two different things. Leadership must be willing to pivot from non-profit to commercialization.
  • The first pitch deck from an academic team is almost always bad. Relying on a personal endorsement from Berkeley professor Scott Shenker—calling Zaharia the best distributed systems scholar of the decade—allowed Horowitz to bet on the people despite the deck. In early-stage investing, personal credibility can override any deck.
  • Technical excellence does not equal a correct business model. Databricks spent years on data lakes and ML before taking off by shifting to the Snowflake warehouse market and rebranding as enterprise AI infrastructure. Every shift in the competitive landscape is a critical re-pricing opportunity.
  • Platform companies commoditize expensive tools (MLflow, Unity Catalog) to force focus back to the primary revenue stream (compute usage). This is a classic platform strategy to protect compute-based revenue moats. By making governance and lifecycle tools free, Databricks effectively offloaded commoditization costs onto competitors.
  • In the AI era, enterprise sensitivity regarding private data is paramount. Running Anthropic Claude, OpenAI GPT, and Google Gemini within Databricks' security perimeter ensures data stays in-domain. This makes data platforms more suitable for enterprise AI infrastructure than pure model companies, shifting the valuation narrative from Big Data SaaS to AI infrastructure.

Core Data

  • Founded:2013 (Company disclosure, as of 2026, independent verification pending)
  • Number of Founders:7 (Company disclosure, as of 2026, independent verification pending)
  • Series A (Sept 2013):$13.9 million, led by a16z (Company disclosure, as of 2026, independent verification pending)
  • Series E (Feb 2019):$250 million, $2.75 billion valuation (Company disclosure, as of 2026, independent verification pending)
  • Series G (Feb 2021):$1 billion, $28 billion valuation, ARR approx. $425 million (Company disclosure, as of 2026, independent verification pending)
  • Series H (Aug 2021):$1.6 billion, $38 billion valuation, ARR over $600 million (Company disclosure, as of 2026, independent verification pending)
  • Series I (Sept 2023):$500 million, $43 billion valuation (Nvidia participated) (Company disclosure, as of 2026, independent verification pending)
  • FY 2023 Revenue:$1.6 billion (Company disclosure, as of 2026, independent verification pending)
  • MosaicML Acquisition (June 2023):$1.3 billion, approx. $21 million per employee (Company disclosure, as of 2026, independent verification pending)
  • Tabular Acquisition (June 2024):Over $1 billion (Company disclosure, as of 2026, independent verification pending)
  • Series J (Dec 2024):$10 billion, $62 billion valuation (led by Thrive Capital) (Company disclosure, as of 2026, independent verification pending)
  • Series K (Aug 2025):$1 billion, over $100 billion valuation, ARR $4 billion (AI revenue over $1 billion run rate) (Company disclosure, as of 2026, independent verification pending)
  • Neon Acquisition (May 2025):Approx. $1 billion (Company disclosure, as of 2026, independent verification pending)
  • Series L (Dec 2025):$4 billion, $134 billion valuation (Company disclosure, as of 2026, independent verification pending)
  • ARR (Feb 2026):$5.4 billion, over 65% YoY growth, free cash flow turned positive for the first time (Company disclosure, as of 2026, independent verification pending)
  • ARR (June 2026):$6.9 billion, 80% YoY growth (Company disclosure, as of 2026, independent verification pending)
  • Series M (July 2026):$3 billion, $188 billion valuation (Company disclosure, as of 2026, independent verification pending)
  • Employees (2026):Approx. 10,000 (Company disclosure, as of 2026, independent verification pending)

Competitors / Peers

The direct competitor is Snowflake. Both companies started from different ends—data lakes (Databricks) and data warehouses (Snowflake)—and collided on the 'Lakehouse' concept. The 2021 benchmark war and Snowflake's adoption of Apache Iceberg to counter Delta Lake merged the markets, leading to a direct fight for total enterprise data budgets. AWS, Google Cloud, and Microsoft Azure are simultaneously shareholders, sales channels, and potential competitors, as each promotes managed Spark or similar ML platforms. In the generative AI infrastructure space, they face direct competition from OpenAI, Anthropic, and Google, but Databricks chooses to integrate their models into its security perimeter, turning rivals into partners while blocking pure model companies that lack a data platform foundation. They also compete indirectly with the Apache Iceberg camp on open-source Lakehouse formats; the 2024 acquisition of Tabular (the Iceberg founder's project) for over $1 billion partially neutralized this format war.