Databricks promises cheap cloud data warehousing

Databricks, the company born out of the Apache Spark boom, has let loose a raft of updates at its San Francisco conference, including an elastic compute option for analytics.
Databricks SQL Serverless, available in preview on AWS, has been designed to improve query performance and concurrency of BI and analytics workloads on messy data lake repositories.
The move is part of the company’s plan to bring data lakes and data warehouses together on one system: the proverbial “lakehouse”, the coinage du jour achieving currency among vendors and commentators alike.
Databricks is also announcing an update to Photon, its query engine for lakehouse systems, making it available in Databricks Workspaces — the environment where users view their Databricks assets. It is releasing Open source connectors for Go, Node.js, and Python to simplify to access a lakehouse from operational applications. Meanwhile, query federation in Databricks SQL is set to let users query data in PostgreSQL, MySQL, AWS Redshift from Databricks.
First announced last year, Databricks SQL Serverless is designed to provide instant compute to users for their BI and SQL workloads. The company promised “minimal management” and capacity optimizations to lower overall cost by an average of 40 per cent.
Joel Minnick, Databricks marketing VP, told The Register that SQL Serverless will allows users coming into Databricks to “go from start up to query in three seconds”.
It also makes economic sense in the cloud, he said. “You truly only pay for what you use on a data warehouse and workloads. That is a real game-changer in terms of the cost of getting these workloads done.”
Minnick said the nearest cloud data warehouse rival would be eight times more expensive than Databricks for these types of workloads.
The serverless option would also help address the issue of user concurrency on analytics queries, where data lakes have attracted criticism.
Minnick said Databricks had already made progress on this issue, and the vast majority of enterprise data warehousing concurrency needs would be met by Databricks SQL, he said.
Hyoun Park, CEO and chief analyst at Amalgam Insights says Databricks’ Serverless SQL makes it easier to support large amounts of distributed data in a cost-effective manner. “It is a response to other vendors providing serverless SQL offerings such as Azure SQL or CockroachDB, but this should also allow Databricks customers to more easily support multi-region and hybrid multi-cloud environments. From a practical perspective, this move makes it easier for potential Databricks customers to use a lakehouse without the significant challenges of manual resource management that can potentially occur as data gets bigger and faster from many different sources to multiple different destinations.


