Apache Storm
Apache Storm is a distributed real-time computation system that processes unbounded streams of data with low latency and high throughput.
Publisher review
Apache Storm is a distributed real-time computation system that processes unbounded streams of data with low latency and high throughput. It is designed for developers and data engineers who need to build reliable stream processing pipelines for use cases such as real-time analytics, fraud detection, and operational monitoring. Storm's architecture centers on a Directed Acyclic Graph (DAG) topology, where spouts ingest data from sources like queues or databases, and bolts perform transformations, aggregations, or other computations. The system guarantees at-least-once processing of every message, supports stateful computations, and integrates with multiple languages including Java, Python, and Ruby. Storm was originally created by Nathan Marz at BackType, which Twitter acquired in 2011; it remains free and open source under the Apache License 2.0.
Storm's key capabilities are built around its topology-based design. A topology defines the flow of data through spouts and bolts, which can be arranged in complex DAGs for multi-stage processing. The system achieves benchmark performance of over one million tuples processed per second per node, making it suitable for high-velocity data streams. Storm provides a simple API for defining topologies, supports both reliable and unreliable spouts to balance performance and durability, and includes built-in fault tolerance that allows the cluster to recover from node failures without data loss. It integrates with common queueing systems like Apache Kafka and databases such as Cassandra, enabling straightforward ingestion and output.
In the stream processing market, Storm faces strong competition from Apache Flink, Apache Kafka Streams, and Apache Spark Streaming. Flink offers more advanced features for event-time processing, memory utilization, and unified stream and batch processing, which has contributed to Storm's declining market share. Kafka Streams provides a lightweight library approach without a separate cluster, while Spark Streaming uses micro-batching for near-real-time processing. Storm's simpler setup and lower operational overhead can be advantageous for teams that prioritize rapid deployment over advanced capabilities, but it lacks the sophisticated event-time semantics and SQL support found in Flink.
The honest trade-offs with Storm involve its maturity and feature set. While it excels at low-latency, high-throughput stream processing with straightforward configuration, it is less suitable for unified batch and stream workloads or complex event processing (CEP) scenarios. Its API is less declarative than Flink or Akka, requiring more manual topology management. Storm's declining community momentum means fewer new features and integrations compared to alternatives, and its reliance on at-least-once semantics may not suit use cases requiring exactly-once guarantees without additional tooling. Organizations with existing Kafka infrastructure may find Kafka Streams simpler, while those needing advanced analytics often prefer Flink.
How it works
-
Distributed real-time computation
Processes unbounded data streams with low latency across a cluster of nodes, enabling real-time analytics and monitoring.
-
Fault-tolerant and reliable
Automatically restarts failed workers and replays tuples to guarantee at-least-once processing without data loss.
-
Scalable and easy to set up
Clusters scale horizontally by adding nodes, with out-of-the-box configurations suitable for production deployment.
-
Integrates with queueing and database technologies
Connects to Apache Kafka, RabbitMQ, Cassandra, and other systems via spouts and bolts for data ingestion and output.
-
Supports multiple languages
Topologies can be written in Java, Python, Ruby, Clojure, or other JVM languages using a multilang protocol.
-
Guarantees data processing
Every tuple emitted by a spout is fully processed or replayed, ensuring no data is lost even during failures.
-
Provides a simple API
Developers define topologies with spouts and bolts using a straightforward Java API, reducing the learning curve.
Strengths and trade-offs
Strengths
- Achieves benchmark performance of over one million tuples processed per second per node for high-throughput workloads.
- Simple setup and configuration process with out-of-the-box settings suitable for production use, reducing deployment time.
- Guarantees at-least-once processing of messages through a reliable acking mechanism, ensuring data integrity.
- Supports both reliable and unreliable spouts, giving developers flexibility to trade durability for lower latency.
Trade-offs
- Market share is declining compared to more advanced alternatives like Apache Flink, which offers better event-time processing and memory utilization.
- Lacks unified stream and batch processing capabilities, making it less suitable for pipelines that require both modes.
- Provides less support for complex event processing (CEP) and SQL on streaming data compared to Flink and Kafka Streams.
- Its API is less declarative than frameworks like Flink or Akka, requiring more manual topology management and configuration.
Pricing context
Free and open source under the Apache License, Version 2.0. No licensing fees; users incur only infrastructure costs for cluster deployment.
Getting started with Apache Storm
-
Download and install Storm
Download the latest Apache Storm release from the official website. Extract the archive to a directory on your cluster nodes. Set the STORM_HOME environment variable and add Storm's bin directory to your PATH.
-
Configure cluster settings
Edit the storm.yaml configuration file to set ZooKeeper connection strings, Nimbus host, and worker ports. Define the number of workers and slots per node according to your cluster resources.
-
Define a topology
Write a Java class that creates a TopologyBuilder object. Add a spout to read from a data source like Kafka, then chain bolts to perform transformations or aggregations. Set parallelism hints for each component.
-
Submit the topology
Use the Storm CLI command 'storm jar' to submit your topology JAR file to the Nimbus server. Specify the main class and any arguments. The cluster will distribute and execute the topology across worker nodes.
-
Monitor and manage topology
Use the Storm UI web interface to view topology statistics, including throughput, latency, and error counts. Kill or rebalance the topology as needed using the CLI commands 'storm kill' or 'storm rebalance'.
Frequently Asked Questions
What is Apache Storm used for?
Apache Storm is a distributed real-time computation system for processing unbounded data streams with low latency and high throughput. It is used for real-time analytics, fraud detection, and operational monitoring by developers and data engineers.
How does Apache Storm's topology work?
A Storm topology is a Directed Acyclic Graph where spouts ingest data from sources like queues, and bolts perform transformations or aggregations. This structure enables multi-stage stream processing with high throughput and low latency across a cluster.
Does Apache Storm guarantee exactly-once processing?
No, Apache Storm guarantees at-least-once processing through a reliable acking mechanism that replays tuples if failures occur. This ensures no data loss but may not suit use cases requiring exactly-once semantics without additional tooling.
How does Apache Storm compare to Apache Flink?
Apache Flink offers more advanced event-time processing, memory utilization, and unified stream and batch capabilities, contributing to Storm's declining market share. Storm has simpler setup and lower operational overhead but lacks Flink's sophisticated features and SQL support.
What are the main strengths of Apache Storm?
Storm achieves over one million tuples processed per second per node, has simple setup with out-of-the-box production configurations, guarantees at-least-once processing, and supports both reliable and unreliable spouts for flexibility between durability and latency.
What are the weaknesses of Apache Storm?
Storm's market share is declining due to competition from Flink and Kafka Streams. It lacks unified batch and stream processing, has less support for complex event processing and SQL, and its API requires more manual topology management than declarative alternatives.
Alternatives
How Apache Storm compares
Direct head-to-head against 2 competitors. Picked by 7wData.
Apache Storm
- Pricing
- Free and open source under the Apache License, Version 2.0. No licensing fees; users incur only infrastructure costs for cluster deployment.
- Target
- Apache Storm is a distributed real-time computation system that processes unbounded streams of data with low latency and high throughput.
- Strength
- Achieves benchmark performance of over one million tuples processed per second per node for high-throughput workloads.
- Watch for
- Market share is declining compared to more advanced alternatives like Apache Flink, which offers better event-time processing and memory utilization.
Apache Flink
- Pricing
- Free (Open Source)
- Target
- Unified stream and batch processing
- Deployment
- On-prem, cloud
- Strength
- Native streaming, event-time processing
- Watch for
- Complex setup for advanced features
Apache Spark Streaming
- Pricing
- Free (Open Source)
- Target
- Micro-batch stream processing
- Deployment
- On-prem, cloud
- Strength
- Integration with Spark ecosystem
- Watch for
- Higher latency compared to Storm
User reviews
No user reviews yet. Be the first to write one.
Sources
Reporting on this tool draws on these publicly available sources.