How does One Consume an Ocean of Data? A Meaningful Sip at a Time

So many data, so many ways to use it, ignore it, misapply it, co-opt, brag, and lament about it. It’s the new oil as suggested not long ago byClive Humby, data scientist, and has been written of recently by authorities such asBernard Marrin Forbeswherein he discusses the apt and not so apt comparison of data and oil. Data are, or data is? Can’t even fully agree on that application of the plural (I’m in the ‘are’ camp.) There’s an ongoing and serious debate on who ‘owns’ data- is possession 9/10 of the law? Not if one considers the regs ofGDPR, and since few industries possess, use, leverage and monetize data more than the insurance industry forward-thinking industry players need to have a well-considered plan for working with data, for, at the end of the day it’s not having the oil, but having the refined byproduct of it, correct?
Tim Stack of technologies solutions company, Cisco, has blogged that 5 quintillion bytes of data are produced daily by IoT devices. That’s 5,000,000,000,000,000,000 bytes of data; if each were a gallon of oil the volume would more than fill the Atlantic Ocean. Just IoT generated bits and bytes. Yes, we have data, we are flush with it. One can’t drink the ocean, but must deal with it, yes?
I was fortunate to be able to broach the topic of data availability with two smart technologists who are also involved with the insurance industry, Lakshan De Silva, CTO of Intellect SEEC, and Christopher Frankland , Head of Strategic Partnerships, ReSource Pro and Founder, InsurTech 360″. Turns out there is so much to discuss that the volume of information would more than fill this column- not by an IoT quintillions’ factor but a by a lot.
With so much data to consider, it’s agreed between the two that understanding the need of data usage guides the pursuit. Machine Learning (ML) is a popular and meaningful application of data, and “can bring with it incredible opportunity around innovation and automation. It is however, indeed a Brave New World,” comments Mr. Frankland. Continuing, “Unless you have a deep grasp or working knowledge of the industry you are targeting and a thorough understanding of the end-to-end process, the risk and potential for hidden technical debt is real.”
What? Too much data, ML methods to help, but now there’s ‘hidden technical debt’ issues? Oil is not that complicated- extract, refine, use. (Of course as Bernard Marr reminds us there are many other concerns with use of natural resources.) Data- plug it into algorithms, get refined ML results. But as noted in Hidden Technical Debt in Machine Learning Systems, ML brings challenges of which data users/analyzers must be aware- compounding of complex issues. ML can’t be allowed to play without adult supervision, else ML will stray from the yard.
From a different perspective Mr. De Silva notes that the explosion of data (and availability of those data) is, “another example of disruption within the insurance industry.” Traditional methods of data use (actuarial practices) are one form of analysis to solve risk problems, but there is now a tradeoff of “what risk you understand upfront”, and “what you will understand through the life of a policy.” Those IoT (or, IoE- Internet of Everything, per Mr. De Silva) data that accumulate in such volume can, if managed/assessed efficiently, open up ‘pay as you go’ insurance products and fraud tool opportunities.
Another caution from Mr. De Silva- assume all data are wrong unless you prove it otherwise. This isn’t as threatening a challenge as it sounds- with the vast quantity and sourcing of data- triangulation methods can be applied to provide a tighter reliability to the data, and (somewhat counterintuitively,) with the analysis of unstructured data with structured across multiple providers and data connectors one can be helped to achieve ‘cleaner’ (reliable) data.


