Why Modern Business Runs On Data Streaming

Data moves. Almost all sources of data have an element of dynamism and motion about them. Even data at rest in some form of archival storage tier will have previously have led a more fluid life moving between applications, devices and network backbones and, inevitably, will have also moved to its resting place via a transport mechanism.
But although almost all data moves, not all data moves at the same speed, at the same cadence and with the same kind of system circularity, size or value.
In the always-on world of cloud computing, mobile device ubiquity and with the planet’s new population of intelligent ‘edge’ machines in the Internet of Things (IoT), a continuous flow of data exists that we naturally refer to as a stream.
So then, what is data streaming and how should we understand it and work with it?
Data streaming is a computing principle and operational system reality where (typically small in size) elements of data travel through an IT system in a time-ordered sequence. Often referred to in the same breath as IT ‘events’ (everything from a user pressing a button on a mouse or a keyboard keystroke… and onward to asynchronous changes that happen as applications execute code and perform their work), the smaller elements of data that form data streaming flows might be made up of log files (small records tracing every behavioral step taken by applications and services), financial transaction logs, web browser activity records, IoT smart-machine sensor readings, geospatial or telemetry information, in-game video game movement and action information… and everything down to the smallest device instrumentation record, which like everything else here creates a data droplet that forms part of a continuous flow that ultimately makes up a stream.
An organization that works with a data-driven data-centric and data-derived approach can take practical steps to analyze its data streaming pipelines in real-time to provide a granular and accurate view of what’s happening in the business. By using a data streaming platform to perform sequentially-processed analysis of every data record in the stream, an organization can sample, filter, correlate and aggregate its data streaming pipeline and start to create a new layer of business insight and control.
Starting simple, a business might choose to build simple real-time streaming alert applications that flag minimum and maximum values to drive alarms and alerts when chosen metrics fall under or exceed pre-specified thresholds. Moving onwards, that same business might then think about applying Machine Learning (ML) algorithms to their data streaming pipeline to look for deeper trends that may be surfacing on a long-term (and eventually, also a shorter-term) basis.
If that’s data streaming for (quite smart) dummies and tech-aware business-tech people covered, then please feel good about ingesting this essential exposition and explanation i.e. this is a technology that is currently being applied to applications in every conceivable business vertical.
The shape of the data streaming market is typical of the enterprise cloud space at large. There are offerings from all of the major cloud services provider hyperscalers (AWS, Google Cloud Platform and Microsoft Azure), IBM has its finger in the pie and there are a group of IT vendors traditionally known for their enterprise data management and integration platforms (Tibco is a good example) that also enjoy share of voice.
Then there is open source, which in this case centralizes around Apache Kafka, an open source data stream processing platform written in the Java and Scala languages. Supporting the use of Kafka at the enterprise level is Confluent.


