How Apache Kafka promises to be your enterprise’s central nervous system for data

Given that all the best big data infrastructure is open source, why do enterprises still spend so heavily? According to new Wikibon research, the big data market will approach $40 billion this year and soar to $100 billion within the next 10 years.
And yet, as Gartner analyst Nick Heudecker captures in a customer complaint, “Why am I paying all these vendors for what’s effectively open source software?”
In the case of Confluent, the company behind the Apache Kafka technology first developed by LinkedIn, the answer is all about packaging. This strategy—build and promote a popular open source project and then monetize management tooling around it—is now a well-trod path, but seems to be particularly fruitful for Confluent.
Though “big data” used to be synonymous with Hadoop, it has come to comprise a host of software—nearly all of it open source—that includes things as varied as MongoDB, Apache Spark, and Apache Kafka. Despite the open source nature of much of this software, there’s a lot of money to be made (Figure A).
The first step in monetizing this open source bounty, however, is popularity. No one will bother to pay for support, much less tooling to make adoption of a particular project more productive, for a random project with minuscule adoption.
This isn’t a problem for Apache Kafka, however.
Apache Kafka is already in production in thousands of companies around the world, including more than one-third of the Fortune 500 and the majority of Silicon Valley‘s tech giants. The reason is simple: Apache Kafka allows companies to go from treating data as something static, that sits in data warehouses or so-called “data lakes,” and enables them to instead build on top of real-time data streams that change continuously with their business.
If this sounds disruptive to the old guard of data infrastructure, it is.


