Apache Storm

Apache Storm is a distributed real-time computation system that processes unbounded streams of data with low latency and high throughput.

Reviewed by 7wData

On this page

Publisher review

Apache Storm is a distributed real-time computation system that processes unbounded streams of data with low latency and high throughput. It is designed for developers and data engineers who need to build reliable stream processing pipelines for use cases such as real-time analytics, fraud detection, and operational monitoring. Storm's architecture centers on a Directed Acyclic Graph (DAG) topology, where spouts ingest data from sources like queues or databases, and bolts perform transformations, aggregations, or other computations. The system guarantees at-least-once processing of every message, supports stateful computations, and integrates with multiple languages including Java, Python, and Ruby. Storm was originally created by Nathan Marz at BackType, which Twitter acquired in 2011; it remains free and open source under the Apache License 2.0.

Storm's key capabilities are built around its topology-based design. A topology defines the flow of data through spouts and bolts, which can be arranged in complex DAGs for multi-stage processing. The system achieves benchmark performance of over one million tuples processed per second per node, making it suitable for high-velocity data streams. Storm provides a simple API for defining topologies, supports both reliable and unreliable spouts to balance performance and durability, and includes built-in fault tolerance that allows the cluster to recover from node failures without data loss. It integrates with common queueing systems like Apache Kafka and databases such as Cassandra, enabling straightforward ingestion and output.

In the stream processing market, Storm faces strong competition from Apache Flink, Apache Kafka Streams, and Apache Spark Streaming. Flink offers more advanced features for event-time processing, memory utilization, and unified stream and batch processing, which has contributed to Storm's declining market share. Kafka Streams provides a lightweight library approach without a separate cluster, while Spark Streaming uses micro-batching for near-real-time processing. Storm's simpler setup and lower operational overhead can be advantageous for teams that prioritize rapid deployment over advanced capabilities, but it lacks the sophisticated event-time semantics and SQL support found in Flink.

The honest trade-offs with Storm involve its maturity and feature set. While it excels at low-latency, high-throughput stream processing with straightforward configuration, it is less suitable for unified batch and stream workloads or complex event processing (CEP) scenarios. Its API is less declarative than Flink or Akka, requiring more manual topology management. Storm's declining community momentum means fewer new features and integrations compared to alternatives, and its reliance on at-least-once semantics may not suit use cases requiring exactly-once guarantees without additional tooling. Organizations with existing Kafka infrastructure may find Kafka Streams simpler, while those needing advanced analytics often prefer Flink.

Get the AI & data signal, daily.

48k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

How it works

  1. Distributed real-time computation

    Processes unbounded data streams with low latency across a cluster of nodes, enabling real-time analytics and monitoring.

  2. Fault-tolerant and reliable

    Automatically restarts failed workers and replays tuples to guarantee at-least-once processing without data loss.

  3. Scalable and easy to set up

    Clusters scale horizontally by adding nodes, with out-of-the-box configurations suitable for production deployment.

  4. Integrates with queueing and database technologies

    Connects to Apache Kafka, RabbitMQ, Cassandra, and other systems via spouts and bolts for data ingestion and output.

  5. Supports multiple languages

    Topologies can be written in Java, Python, Ruby, Clojure, or other JVM languages using a multilang protocol.

  6. Guarantees data processing

    Every tuple emitted by a spout is fully processed or replayed, ensuring no data is lost even during failures.

  7. Provides a simple API

    Developers define topologies with spouts and bolts using a straightforward Java API, reducing the learning curve.

Strengths and trade-offs

Strengths

  • Achieves benchmark performance of over one million tuples processed per second per node for high-throughput workloads.
  • Simple setup and configuration process with out-of-the-box settings suitable for production use, reducing deployment time.
  • Guarantees at-least-once processing of messages through a reliable acking mechanism, ensuring data integrity.
  • Supports both reliable and unreliable spouts, giving developers flexibility to trade durability for lower latency.

Trade-offs

  • Market share is declining compared to more advanced alternatives like Apache Flink, which offers better event-time processing and memory utilization.
  • Lacks unified stream and batch processing capabilities, making it less suitable for pipelines that require both modes.
  • Provides less support for complex event processing (CEP) and SQL on streaming data compared to Flink and Kafka Streams.
  • Its API is less declarative than frameworks like Flink or Akka, requiring more manual topology management and configuration.

Pricing context

Free and open source under the Apache License, Version 2.0. No licensing fees; users incur only infrastructure costs for cluster deployment.

Getting started with Apache Storm

  1. Download and install Storm

    Download the latest Apache Storm release from the official website. Extract the archive to a directory on your cluster nodes. Set the STORM_HOME environment variable and add Storm's bin directory to your PATH.

  2. Configure cluster settings

    Edit the storm.yaml configuration file to set ZooKeeper connection strings, Nimbus host, and worker ports. Define the number of workers and slots per node according to your cluster resources.

  3. Define a topology

    Write a Java class that creates a TopologyBuilder object. Add a spout to read from a data source like Kafka, then chain bolts to perform transformations or aggregations. Set parallelism hints for each component.

  4. Submit the topology

    Use the Storm CLI command 'storm jar' to submit your topology JAR file to the Nimbus server. Specify the main class and any arguments. The cluster will distribute and execute the topology across worker nodes.

  5. Monitor and manage topology

    Use the Storm UI web interface to view topology statistics, including throughput, latency, and error counts. Kill or rebalance the topology as needed using the CLI commands 'storm kill' or 'storm rebalance'.

Frequently Asked Questions

What is Apache Storm used for?

Apache Storm is a distributed real-time computation system for processing unbounded data streams with low latency and high throughput. It is used for real-time analytics, fraud detection, and operational monitoring by developers and data engineers.

How does Apache Storm's topology work?

A Storm topology is a Directed Acyclic Graph where spouts ingest data from sources like queues, and bolts perform transformations or aggregations. This structure enables multi-stage stream processing with high throughput and low latency across a cluster.

Does Apache Storm guarantee exactly-once processing?

No, Apache Storm guarantees at-least-once processing through a reliable acking mechanism that replays tuples if failures occur. This ensures no data loss but may not suit use cases requiring exactly-once semantics without additional tooling.

How does Apache Storm compare to Apache Flink?

Apache Flink offers more advanced event-time processing, memory utilization, and unified stream and batch capabilities, contributing to Storm's declining market share. Storm has simpler setup and lower operational overhead but lacks Flink's sophisticated features and SQL support.

What are the main strengths of Apache Storm?

Storm achieves over one million tuples processed per second per node, has simple setup with out-of-the-box production configurations, guarantees at-least-once processing, and supports both reliable and unreliable spouts for flexibility between durability and latency.

What are the weaknesses of Apache Storm?

Storm's market share is declining due to competition from Flink and Kafka Streams. It lacks unified batch and stream processing, has less support for complex event processing and SQL, and its API requires more manual topology management than declarative alternatives.

Alternatives

How Apache Storm compares

Direct head-to-head against 2 competitors. Picked by 7wData.

This tool

Apache Storm

Pricing
Free and open source under the Apache License, Version 2.0. No licensing fees; users incur only infrastructure costs for cluster deployment.
Target
Apache Storm is a distributed real-time computation system that processes unbounded streams of data with low latency and high throughput.
Strength
Achieves benchmark performance of over one million tuples processed per second per node for high-throughput workloads.
Watch for
Market share is declining compared to more advanced alternatives like Apache Flink, which offers better event-time processing and memory utilization.

Apache Flink

Pricing
Free (Open Source)
Target
Unified stream and batch processing
Deployment
On-prem, cloud
Strength
Native streaming, event-time processing
Watch for
Complex setup for advanced features

Apache Spark Streaming

Pricing
Free (Open Source)
Target
Micro-batch stream processing
Deployment
On-prem, cloud
Strength
Integration with Spark ecosystem
Watch for
Higher latency compared to Storm

User reviews

No user reviews yet. Be the first to write one.

Sources

Reporting on this tool draws on these publicly available sources.

  1. storm.apache.org
  2. storm.apache.org
  3. storm.apache.org
  4. www.conduktor.io
  5. www.onehouse.ai
  6. taogang.medium.com