Starburst targets data discoverability with new capabilities

Starburst on Wednesday unveiled a trio of new features for its data management and analytics platform, including an automated data catalog for Starburst Galaxy that aims to enable users to more quickly and easily search and discover data.
Galaxy is the vendor’s cloud-native offering, while Enterprise is for its on-premises customers.
In addition to the automated data catalog, the vendor revealed that users can now use Python to access both Galaxy and Enterprise. It also revealed a tool named Warp Speed that automates indexing and caching of data an accelerates queries up to seven times over manual indexing and caching, according to the vendor.
Starburst, founded in 2017 and based in Boston, is a data and analytics vendor whose platform is designed to let customers build a data mesh architecture.
Data mesh is a decentralized approach to data management and analytics that relies on the domain expertise of power users within departments to help oversee their organizations’ data operations. Other vendors offering data mesh tools include Informatica and Talend.
Starburst unveiled its new capabilities during Datanova, a virtual user conference hosted by the vendor.
The automated data catalog and Warp Speed are now in private preview, while access via Python is generally available. Warp Speed is expected to be generally available to Enterprise users by the end of February and Galaxy users within the next three months.
Among the features unveiled during Datanova, the automated data catalog and Warp Speed have the potential to be of most benefit to customers, according to Doug Henschen, an analyst at Constellation Research. Even before developing the automated data catalog, Starburst’s query engine automatically collected metadata about user behavior related to a given dataset when that dataset was connected to Starburst. Now as soon as a dataset is connected to Starburst, the platform not only collects metadata but also adds the dataset to a catalog so it can be found and queried for relevant analytics use. Warp Speed, meanwhile, represents Starburst repackaging of capabilities it inherited through its June 2022 acquisition of Varada. The tool indexes and caches workloads, automatically assembling workloads into blocks based on patterns of use and other characteristics that make it reasonable to group workloads together. Ultimately, Starburst built both tools to make data more discoverable and speed up the process of organizing data. The desired result is to make it faster and easier to find relevant data that leads to insights and actions. “The features that will have the broadest appeal [will] be the indexing and caching feature, which will drive performance, and the automated data catalog, which will make it easier for any users to find data and assets that could be potentially helpful,” Henschen said. “More sophisticated users might develop data products … but performance gains and improved data access benefit every user.” Similarly, Vishal Singh, head of data products at Starburst, cited the automated data catalog as having the most potential significance for users among the new capabilities unveiled during Datanova. Singh noted that it enables users to know which datasets to query without having to first search for them on the different databases an organization might use. It also avoids potentially finding various versions of the same dataset in a few different places and having to explore those datasets to discover which is most relevant.


