The time is now for the synthetic data revolution

3 min read
Curated from itproportal.com →

We’re going into the final sprint of the year. Time will fly by as we get stuck into the festive season, and before we know it, we’ll have a brand new year stretching out ahead of us. 

Not that I like to wish time away, but I can’t wait for 2022. That’s because it looks set to be a huge year for synthetic data. Forbes picked it in a list of the 5 biggest data science trends for 2022, while Gartner put synthetic data at number one in its top strategic predictions for 2022 and beyond. 

If we’re talking about the passing of time, going from founding Mindtech four years ago to the brink of a synthetic data revolution feels like it’s happened in the blink of an eye. 

It seems fitting though, given that time, or time-saving is one of the driving forces behind synthetic data development. Did you know, it takes a long and laborious 20 weeks to gather and annotate the 100,000 real-world images required to train a visual AI system to see and understand the world as a human? 

That’s roughly 80 percent of machine learning project time, just for something novel, like training a system to pick out a lost child in a busy shopping mall. It takes even longer to help a delivery robot service safely navigate spaces where children are playing, leaving no room for network development or gleaning insights from the data.

It’s time for a sea change in the way we view data and how we train AI. Synthetic data derived from computer-generated images and video is easily as good—sometimes better—than that which comes from real-world images and it can rapidly shrink the process of gathering and analyzing from months to hours. 

All this, without impacting on the AIs being trained. That’s because to an AI, there is no ‘real’ or ‘synthetic’; there’s only the data we give it. 

The technology is ready; it’s time for us humans to stop seeing synthetic data as secondary, and start to understand the opportunity in our hands to scale the AI industry exponentially.

It doesn’t matter if they’re start-ups, scale-ups, or global enterprises, teams trying to secure the required high-quality images to train their new AI system will all be up against the ‘Big Four’: Apple, Amazon, Facebook, and Google. Engineers at the latter have access to more than 4 trillion images alone stored in Google Photos. 

These major players tend to restrict access to this wealth of potential training data because it hands them a competitive advantage to develop new products and monetize their datasets. Even they’re not totally immune to the issues the industry faces though; searching through trillions of images to find the relevant ones is non-trivial, and once found, they still need annotating. 

Every company has to navigate the challenge of more readily enforced data privacy regulations too—including the EU’s General Data Protection Regulation (GDPR). Just ask Facebook/Meta, who recently announced it will delete its facial recognition system and database, saying it is doing so because regulators can’t keep up with announcements.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at itproportal.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.