What China’s Algorithm Registry Reveals about AI Governance

3 min read

For the past year, the Chinese government has been conducting some of the earliest experiments in building regulatory tools to govern artificial intelligence (AI). In that process, China is trying to tackle a problem that will soon face governments around the world: Can regulators gain meaningful insight into the functioning of algorithms, and ensure they perform within acceptable bounds?

One particular tool deserves attention both for its impact within China, and for the lessons technologists and policymakers in other countries can draw from it: a mandatory registration system created by China’s internet regulator for recommendation algorithms.

Although the full details of the registry are not public, by digging into its online instruction manual, we can reveal new insights into China’s emerging regulatory architecture for algorithms.

The algorithm registry was created by China’s 2022 regulation on recommendation algorithms (English translation), which came into effect in March of this year and was led by the Cyberspace Administration of China (CAC). China’s algorithm regulation has largely focused on the role recommendation algorithms play in disseminating information, requiring providers to ensure that they don’t “endanger national security or the social public interest” and to “give an explanation” when they harm the legitimate interests of users. Other provisions sought to address monopolistic behavior by platforms and hot-button social issues, such as the role that dispatching algorithms play in creating dangerous labor conditions for Chinese delivery drivers.

The regulation also requires recommendation algorithms with “public opinion characteristics” and “social mobilization capabilities” to complete a filing with the mysteriously named Internet Information Service Algorithm Filing System. But the provisions didn’t elaborate on the specifics, and the filing requirement went largely unremarked upon at the time.

Analysts got a first look at those filings in August 2022, when the CAC released the first batch of thirty algorithm registrations. Accessible via the registry’s web page, the filings included algorithms from some of China’s biggest internet platform companies, including Tencent, Alibaba, and Bytedance. These publicly available filings usually consisted of a single page with six different short-response categories, including “Algorithm Fundamentals” and “Algorithm Operating Mechanism.”

But the actual descriptions in these filings were pitched at such a high level as to be almost completely devoid of meaningful detail. For example, the filing for Weibo’s “hot search” feature describes the algorithm as adding together “search popularity, discussion popularity, and dissemination popularity,” multiplied by an “interaction rate coefficient.” That may be an accurate description, but it is also so high level that an observer with no knowledge of this specific algorithm could essentially guess it. If this were the full extent of information given to Chinese regulators, it would provide them with no meaningful insights into the algorithms, how they were trained, or how they might perform.

But a closer look at the algorithm registry’s landing page, which contains a downloadable user manual for entities registering their algorithms, provides a wider window into what information the CAC is actually gathering. A close read of that manual and examination of the screenshots within it show that only a portion of the information filed in the registry has been revealed.

The most detailed requirements for disclosure in the manual come via a screenshot showing the page for disclosing “Detailed Algorithm Attribute Information.” Here it asks that algorithm providers list the name of each open-source and self-built data set that was used to train the model, as well as the specific source of that data. In addition, it requires the provider to state whether algorithm inputs involve biometric or other personal information.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at carnegieendowment.org →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.