Lakestream
Lakestream is a new architectural paradigm from StreamNative that unifies real-time streaming and the lakehouse by making streams first-class lakehouse primitives.
Publisher review
Lakestream is a new architectural paradigm from StreamNative that unifies real-time streaming and the lakehouse by making streams first-class lakehouse primitives. Founded in 2018 by the creators of Apache Pulsar and headquartered in San Francisco, Lakestream targets data engineers and architects who need to run Kafka or Pulsar workloads while directly storing data in open table formats like Apache Iceberg and Delta Lake, eliminating the need for separate connectors, ETL pipelines, or materialization jobs. It is designed for organizations that want to reduce total cost of ownership (TCO) for streaming infrastructure while enabling analytics, machine learning, and AI agents on the same data without duplication.
The architecture rests on three independent layers: cloud-native stream storage that writes directly to object storage with low-latency writes and Parquet compaction; a unified Lakestream Catalog that federates with Databricks Unity Catalog, Snowflake Horizon Catalog, and AWS S3 Tables; and stateless protocol servers for Kafka, Pulsar, REST, and gRPC. The first proof point is Ursa For Kafka (UFK), a native Apache Kafka 4.2+ fork that delivers up to 95% cost reduction compared to traditional Kafka deployments. UFK is currently in Limited Public Preview and allows users to run Kafka topics and Iceberg tables as the same object — no movement, no connectors, no waiting.
Lakestream competes directly with Apache Kafka, Apache Pulsar, RabbitMQ, and the Solace Platform by pushing interoperability down to the storage and catalog layers rather than the protocol layer. While traditional streaming systems require bridging separate systems, Lakestream treats streaming and the lakehouse as one system. This positions it against lakehouse-native streaming features from Databricks and Snowflake, as well as connector-first vendors like Confluent, but with the advantage that the protocol becomes a choice of interface, not a data silo.
Honest trade-offs include significant feature bloat that can make the software more complicated and fragile than simpler alternatives. Achieving exactly-once message processing remains difficult in practice. New users face a steep learning curve, especially when integrating with external dependencies like Unity Catalog or S3 Tables, which can complicate issue resolution. The platform's reliance on object storage for low-latency writes may introduce variability in performance compared to local-disk-based systems like Redpanda.
How it works
-
Native Kafka & Pulsar
Run native Apache Kafka and Pulsar workloads with zero application changes, using stateless protocol servers that support Kafka, Pulsar, REST, and gRPC.
-
Ursa For Kafka (UFK)
A native Apache Kafka 4.2+ fork that delivers up to 95% cost reduction by leveraging lakehouse-native storage and eliminating separate streaming clusters.
-
Lakehouse-Native Storage
Streams write directly to object storage in open formats like Iceberg and Delta Lake, with low-latency writes and Parquet compaction for analytics and ML.
-
Unified Metadata Catalog
The Lakestream Catalog federates with Databricks Unity Catalog, Snowflake Horizon Catalog, and AWS S3 Tables for a single source of truth.
-
Three-Layer Architecture
Separates data, metadata, and protocol layers, allowing independent scaling of storage, catalog federation, and protocol support.
-
Agent-Ready Streaming
Provides live context from streams with built-in governance, enabling AI agents to consume real-time data without additional pipelines.
-
Real-Time Backbone
Acts as both a streaming backbone and agent runtime, supporting high-throughput, low-latency data delivery across applications.
Strengths and trade-offs
Strengths
- Up to 95% cost reduction with Ursa For Kafka compared to traditional Kafka deployments, as claimed in the April 2026 announcement.
- Unified architecture eliminates connectors and ETL pipelines by making Kafka topics and Iceberg tables the same object.
- Stateless protocol servers support Kafka, Pulsar, REST, and gRPC from a single storage layer, reducing operational complexity.
- Founded by the creators of Apache Pulsar, with seven years of experience in compute-storage separation and lakehouse-native storage.
Trade-offs
- Feature bloat from combining streaming and lakehouse capabilities can make the software more complicated and fragile than dedicated alternatives.
- Achieving exactly-once message processing remains difficult in practice, especially when federating with external catalogs.
- Steep learning curve for new users due to the three-layer architecture and dependencies on object storage and catalog services.
- External dependencies on Databricks Unity Catalog, Snowflake Horizon, or AWS S3 Tables can complicate issue resolution and add vendor lock-in risk.
Pricing context
Subscription-based with tiered plans based on usage, features, and enterprise requirements. Free and paid plans available. Ursa For Kafka is currently in Limited Public Preview with no public pricing details.
Getting started with Lakestream
-
Sign up for Lakestream
Go to the StreamNative website and create an account. Choose a free or paid plan based on your usage needs. Complete the registration process and verify your email to activate your account.
-
Connect object storage
In the Lakestream console, add a new storage connection. Select your cloud provider (AWS, Azure, or GCP) and provide credentials for your S3 or equivalent bucket. Configure the bucket path where streams will write data.
-
Configure a Kafka topic
Create a new Kafka topic in the Lakestream interface. Set the topic name, partition count, and replication factor. Enable lakehouse-native storage by selecting Iceberg or Delta Lake as the output format.
-
Produce and consume messages
Use a Kafka client (e.g., kafka-console-producer) to send messages to your topic. Then run a consumer to read messages back. Verify that data appears in your object storage as Parquet files in the specified format.
-
Federate with a catalog
In the Lakestream Catalog settings, add a federation target such as Databricks Unity Catalog. Provide the catalog endpoint and authentication token. Test the connection to ensure metadata syncs between systems.
Frequently Asked Questions
What is Lakestream and how does it unify streaming and lakehouse?
Lakestream is a StreamNative architecture that makes streams first-class lakehouse primitives. It stores data directly in open formats like Apache Iceberg and Delta Lake, eliminating separate connectors or ETL pipelines. This allows real-time streaming and analytics on the same data without duplication.
How does Ursa For Kafka reduce costs compared to traditional Kafka?
Ursa For Kafka is a native Apache Kafka 4.2+ fork from Lakestream that claims up to 95% cost reduction. It achieves this by leveraging lakehouse-native storage, eliminating separate streaming clusters, and allowing Kafka topics and Iceberg tables to exist as the same object.
What are the main components of Lakestream's three-layer architecture?
Lakestream's architecture has three independent layers: cloud-native stream storage writing to object storage, a unified Lakestream Catalog federating with Databricks Unity Catalog and others, and stateless protocol servers supporting Kafka, Pulsar, REST, and gRPC. This allows independent scaling of each layer.
How does Lakestream integrate with existing data catalogs like Unity Catalog?
Lakestream includes a unified Lakestream Catalog that federates with Databricks Unity Catalog, Snowflake Horizon Catalog, and AWS S3 Tables. This provides a single source of truth for metadata, enabling seamless access to streaming data across different lakehouse environments without additional connectors.
What are the main trade-offs of using Lakestream for streaming?
Lakestream's feature bloat can make it more complicated and fragile than simpler alternatives. Achieving exactly-once message processing remains difficult, and new users face a steep learning curve due to the three-layer architecture and dependencies on external catalogs like Unity Catalog or S3 Tables.
Is Ursa For Kafka available now and what is its pricing?
Ursa For Kafka is currently in Limited Public Preview with no public pricing details. Lakestream offers subscription-based tiered plans based on usage and features, with both free and paid options available. The preview allows users to test the cost reduction and lakehouse-native storage capabilities.
Alternatives
How Lakestream compares
Direct head-to-head against 3 competitors. Picked by 7wData.
Lakestream
- Pricing
- Subscription-based with tiered plans based on usage, features, and enterprise requirements. Free and paid plans available. Ursa For Kafka is currently in Limited Public Preview with no public pricing details.
- Target
- Lakestream is a new architectural paradigm from StreamNative that unifies real-time streaming and the lakehouse by making streams first-class lakehouse primitives.
- Strength
- Up to 95% cost reduction with Ursa For Kafka compared to traditional Kafka deployments, as claimed in the April 2026 announcement.
- Watch for
- Feature bloat from combining streaming and lakehouse capabilities can make the software more complicated and fragile than dedicated alternatives.
StreamNative
- Pricing
- Custom/Contact sales
- Target
- High-performance Kafka workloads
- Deployment
- Cloud
- Strength
- Leaderless architecture reduces costs
- Watch for
- Complex setup for non-Kafka users
Redpanda
- Pricing
- Custom/Contact sales
- Target
- Real-time data streaming
- Deployment
- Cloud, On-prem
- Strength
- High throughput with low latency
- Watch for
- Higher infrastructure costs
Confluent WarpStream
- Pricing
- Custom/Contact sales
- Target
- Enterprise Kafka workloads
- Deployment
- Cloud
- Strength
- Managed Kafka service
- Watch for
- Higher TCO for high-throughput workloads
User reviews
No user reviews yet. Be the first to write one.
Sources
Reporting on this tool draws on these publicly available sources.