Data lakes going the way of the visual spreadsheet?

Self-service analytics comes in different shapes and sizes, and so data lakes. Both are widely popular concepts that have been shaping the big data world, so it’s no wonder that a flurry of approaches and tools exist there.
There is also a fair amount of overlap between the two. Hadoop-based data lakes are rather common these days, but that does not make them easy to work with for the non-data science types. So self-service analytics tools make a point in trying to support them as data sources their users can connect to.
This happens through a layer of mediation, typically SQL-based. There are various SQL-on-Hadoop engines around, ranging from proprietary to open source, and each distribution comes with its own.
So, depending on how fast your SQL-on-Hadoop engine is and how big your big data lake is, your mileage on the self-service tool side will vary. Typically, such tools also try to facilitate things on their side, by supporting as many engines as possible, applying smart connection techniques, and so on.
In any case, the whole point in self-service analytics, as opposed to traditional data warehouses, is to skip the data mediation process. This requires things such as dimension definition and data cube preparation and therefore a team of people whose job is to work on that.
The idea in self-service analytics is to let users explore data sources on their own, on the fly and using visual paradigms. There is a wide range of tools in that category, each with its own approach and strengths, and then there are some deviants, too.
Datameer is one of those deviants. Its paradigm for exploration is the spreadsheet. You may argue the point in using visual tools is to avoid having to go through endless rows and columns, and that this prospect only gets more scary when you are dealing with data at that scale.
However, there obviously is a market segment for which this paradigm is useful. Spreadsheets have been around for a long time, and many people have spreadsheet skills. In essence, Datameer’s platform gives them a way to not stray too much out of their comfort zone, while offering an alternative to SQL-on-Hadoop.
Datameer lets users connect to a variety of Hadoop distributions on premises or in the cloud, and provides a mechanism for entering declarative spreadsheet formulas that are translated to fully optimized Hadoop jobs.
Datameer also supports ETL and visualization features, and you can export your Datameer spreadsheets to work with CSV, Apache AVRO, Parquet and Tableau formats. Now Datameer is adding another feature in its arsenal called visual exploration.
This is an interesting move in keeping with the times. It does not give up on the spreadsheet paradigm, but it gives users the ability to visually go through charts summarizing their endless rows and columns.
Users can choose which fields from their datasets they are interested in, and the Visual Explorer will summarize them in charts, offering the option to drill down as well. Then users can decide whether that’s an interesting slice of their data for further analysis.

