What Happens When AI Has Read Everything?

4 min read
Curated from theatlantic.com →

Artificial intelligence has in recent years proved itself to be a quick study, although it is being educated in a manner that would shame the most brutal headmaster. Locked into airtight Borgesian libraries for months with no bathroom breaks or sleep, AIs are told not to emerge until they’ve finished a self-paced speed course in human culture. On the syllabus: a decent fraction of all the surviving text that we have ever produced.

When AIs surface from these epic study sessions, they possess astonishing new abilities. People with the most linguistically supple minds—hyperpolyglots—can reliably flip back and forth between a dozen languages; AIs can now translate between more than 100 in real time. They can churn out pastiche in a range of literary styles and write passable rhyming poetry. DeepMind’s Ithaca AI can glance at Greek letters etched into marble and guess the text that was chiseled off by vandals thousands of years ago.

These successes suggest a promising way forward for AI’s development: Just shovel ever-larger amounts of human-created text into its maw, and wait for wondrous new skills to manifest. With enough data, this approach could perhaps even yield a more fluid intelligence, or a humanlike artificial mind akin to those that haunt nearly all of our mythologies of the future.

The trouble is that, like other high-end human cultural products, good prose ranks among the most difficult things to produce in the known universe. It is not in infinite supply, and for AI, not any old text will do: Large language models trained on books are much better writers than those trained on huge batches of social-media posts. (It’s best not to think about one’s Twitter habit in this context.) When we calculate how many well-constructed sentences remain for AI to ingest, the numbers aren’t encouraging. A team of researchers led by Pablo Villalobos at Epoch AI recently predicted that programs such as the eerily impressive ChatGPT will run out of high-quality reading material by 2027. Without new text to train on, AI’s recent hot streak could come to a premature end.

It should be noted that only a slim fraction of humanity’s total linguistic creativity is available for reading. More than 100,000 years have passed since radically creative Africans transcended the emotive grunts of our animal ancestors and began externalizing their thoughts into extensive systems of sounds. Every notion expressed in those protolanguages—and many languages that followed—is likely lost for all time, although it gives me pleasure to imagine that a few of their words are still with us. After all, some English words have a shockingly ancient vintage: Flow, mother, fire, and ash come down to us from Ice Age peoples.

Writing has allowed human beings to capture and store a great many more of our words. But like most new technologies, writing was expensive at first, which is why it was initially used primarily for accounting. It took time to bake and dampen clay for your stylus, to cut papyrus into strips fit to be latticed, to house and feed the monks who inked calligraphy onto vellum. These resource-intensive techniques could preserve only a small sampling of humanity’s cultural output.

Not until the printing press began machine-gunning books into the world did our collective textual memory achieve industrial scale. Researchers at Google Books estimate that since Gutenberg, humans have published more than 125 million titles, collecting laws, poems, myths, essays, histories, treatises, and novels. The Epoch team estimates that 10 million to 30 million of these books have already been digitized, giving AIs a reading feast of hundreds of billions of, if not more than a trillion, words.

Those numbers may sound impressive, but they’re within range of the 500 billion words that trained the model that powers ChatGPT. Its successor, GPT-4, might be trained on tens of trillions of words.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at theatlantic.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.