The pros and cons of data integration architectures

3 min read
Curated from itproportal.com →

The first step in any large-scale data integration project provides businesses and government departments alike with a lot of architectural choices to make when it comes to building new applications. Whether those applications are operational or analytical, the architectural choices that accompany them come with a long list of pros and cons.

While the perceived wisdom might be to build a solution from a variety of different components to increase flexibility and potentially lower costs, there is still a need to integrate all of these various elements together and enforce the necessary levels of governance and security.

The challenge is deciding what the best approach for this is. And this challenge, whether or not it’s faced by an enterprise or public sector department, is made even harder when you consider the volume of legacy data likely to exist across the organisation, which all too often is stored in silos.

Much of the data owned and stored by businesses and government departments alike is constrained by the silos it’s stuck in, many of which have been built over the years as organisations grow. When you consider the consolidation of both legacy and new IT systems, the number of these data silos only increases. What’s more, the impact of this is significant. It has been widely reported that up to 80 per cent of a data scientist’s time is spent on collecting, labelling, cleaning and organising data in order to get it into a usable form for analysis.

Successfully integrating data in order to interrogate it when faced with this dilemma is therefore not an easy process, but there are a number of options available to tackle this head on. However, as mentioned, it’s not necessarily immediately obvious which of these options is best. Data Hub? Data Lake? Virtual Database? What is clear though is that security is a paramount requirement. Unless security is addressed as a core principal rather than an afterthought, businesses and public sector departments leave themselves wide open for things to fall through the cracks, particularly when it comes to open source.

In fact, we saw evidence of what can happen when things go wrong just last summer, when a significant breach was identified by security researchers in a biometrics system widely used by banks as well as defence contractors and the Metropolitan Police. According to reports, the database used to store the facial recognition information and fingerprints of over one million people was discovered to be unprotected and largely unencrypted. This meant that the researchers had access to some 23GB of data, which reportedly also included security levels and clearance, as well as personal details of staff.

This goes to show that when it comes to highly sensitive information such as biometric data, which is increasingly being used by public sector authorities as well as private sector organisations, you cannot afford to take any short cuts. In so many of the publicly reported cases of data breaches, the platforms involved rely entirely upon network security alone.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at itproportal.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.