Privacy by Design laws will kill your data pipelines

3 min read
Curated from protocol.com →

The legislation could make old data pipelines more trouble than they’re worth.
Data pipelines have become so unwieldy that companies might not even know if they are complying with regulations.

A car is totaled when the cost to repair it exceeds its total value. By that logic, Privacy by Design legislation could soon be totaling data pipelines at some of the most powerful tech companies.

Those pipelines were developed well before the advent of more robust user privacy laws, such as the European Union’s GDPR ( 2018 ) and the California Consumer Privacy Act ( 2020 ). Their foundational architectures were therefore designed without certain privacy-preserving principals in mind, including k-anonymity and differential privacy.

But the problem extends way beyond trying to layer privacy mechanisms on top of existing algorithms. Data pipelines have become so complex and unwieldy that companies might not even know whether they are complying with regulations. As Meta engineers put it in a leaked internal document : “We do not have an adequate level of control and explainability over how our systems use data, and thus we can’t confidently make controlled policy changes or external commitments.”

(When we asked Meta for comment, a spokesperson referred us to the company’s original response to Motherboard about the leaked document, which said, in part: “The document was never intended to capture all of the processes we have in place to comply with privacy regulations around the world or to fully represent how our data practices and controls work.”)

As governments increasingly embrace Privacy by Design (PbD) legislation, tech companies face a choice: either start from scratch or try to fix data pipelines that are old, extraordinarily complex and already non-compliant. Some computer science researchers say a fresh start is the only way to go. But for tech companies, starting over would require engineers to roll out critical data infrastructure changes without disrupting day-to-day operations — a task that’s easier said than done.

‘Open borders’ won’t cut it

Motherboard published the leaked internal document, written by Meta engineers in 2021, at the end of April. In it, an engineering team recommended data architecture changes that would help Meta comply with a wave of governments embracing the “consent regime,” one of the core principles of PbD. India, Thailand, South Korea, South Africa and Egypt were all preparing “impactful regulations” in this realm, and the paper also anticipated U.S. federal privacy regulation in 2022 and beyond. Such legislation would generally require Meta to obtain user consent before collecting data for advertisements.

The Meta engineers identified “the heart of our challenge” as a lack of “closed form systems.” Closed systems, they said, would let Meta enumerate and control all the incoming data flows. The engineers placed that in contrast with the “open borders” system that had been baked into company culture for over a decade.
Meta’s systems had grown increasingly complex and untraceable, the engineers said, citing the example of a single feature (“user_home_city_moved”) drawing from around six thousand data tables.

“These are massive pipelines with massive amounts of data feeding into many different kinds of algorithms,” Nikola Banovic, an assistant professor of computer science and engineering at the University of Michigan, told Protocol. “Because this was never a consideration to begin with, now it’s increasingly difficult to untangle things.”

The leaked document showed the frustration of internal teams tasked with overhauling systems designed in an era when everything was fair game, Banovic said. He noted that advocacy groups are pressuring companies to now design systems around end users.

“It’s not going to be easy,” Banovic said of the shift.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at protocol.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.