Synthetic data is the renewable source we need to accelerate the AI industry

It takes an astonishing 20 weeks to gather and annotate the 100,000 real-world images necessary to train a visual AI system to see and understand the world as humans
If you started annotating today, your new year’s resolutions will have comfortably been made and broken by the time you finish.
And that’s just for something novel, like training a system to pick out a lost child in a busy shopping mall. It takes even more images to help a delivery robot service safely navigate spaces where children are playing.
The data scientists working on these systems can spend up to 80% of their time gathering, cleaning, and manually annotating real-world images to be digested by AI systems. It’s too long. It doesn’t leave any time for network development or gleaning insights from the data.
What happens if a project needs the system to fire faster? Or there’s a shortage of images? We’ve recently seen in the UK the real-world consequences of what happens when a scarce resource runs low. Everything grinds to a halt.
Thankfully, there is an alternative, ethical, and endlessly renewable material that can be used to train AI: synthetic data derived from computer-generated images and video. Data produced in this way is easily as good, sometimes even better quality, than that which comes from real-world images, and working with it cuts down the labour intensive process of gathering and analysing images from months to hours.
Crucially, switching has no impact on the AIs being trained. To an AI, there is no ‘real’ or ‘synthetic’; there’s only the data we give it. It’s us humans that need to stop seeing synthetic as some ersatz alternative, and start to understand the opportunity in our hands to scale the AI industry exponentially—if engineers are willing to embrace the synthetic data route.
It doesn’t matter if they’re start-ups, scale-ups, or even a global enterprise company, teams trying to access enough high-quality images to train their new AI system are competing against the ‘Big Four’: Apple, Amazon, Facebook, and Google. Engineers at the search engine giant have access to more than 4 trillion images alone stored in Google Photos.
These major players tend to restrict access to this wealth of potential training data because it secures their competitive advantage to develop new products and monetise their datasets. They’re not totally immune to the issues the rest of the industry faces though; searching through trillions of images to find the relevant ones is non-trivial, and once found, they still need annotating.
Every company has to navigate the challenge of more readily enforced data privacy regulations too—including the EU’s General Data Protection Regulation (GDPR). Just ask Microsoft, who in 2019 deleted its database of 10 million images—the largest publicly available facial recognition data set at the time—due to data privacy concerns.
These factors combine to create that scarcity of real-world visual data I was talking about, with only the very largest tech companies able to compete, driving down the competition and, ultimately, quality of AI systems on the market.
If we want the best AI, the best technology, then we need a competitive landscape made of businesses of all sizes pushing each other forward.


