New algorithm can create movies from just a few snippets of text

2 min read
Curated from sciencemag.org →

Screenwriters denied the big budgets and formidable resources of the major film studios may soon have another option, thanks to a new algorithm that can generate a video simply by consuming a (very short) script. The new movies are far from Oscar-worthy, but a similar technique could one day find uses outside entertainment, by, say, helping a witness reconstruct a car crash or a crime.

Artificial intelligence (AI) is getting much better at identifying the content of images and providing labels. So-called “generative” algorithms go the other way, producing images from labels (or brain scans). A few can even take a single movie frame and predict the next series of frames. But putting it all together—creating an image from text and making it move realistically in accordance with the text—has not been done before.

“As far as I know, it’s the first text-to-video work that gives such good results. They are not perfect, but at least they start to look like real videos,” says Tinne Tuytelaars, a computer scientist at Katholieke Universiteit Leuven in Belgium, who has done her own video prediction research. “It’s really nice work.”

The new algorithm is a form of machine learning, which means it requires training. Specifically, it’s a neural network, or a series of layers of small computing elements that process data in a way reminiscent of the brain’s neurons. During training, software assesses its performance after each attempt, and feedback circulates through the millions of network connections to refine future computations.

This network operates in two stages “designed to mimic how humans create art,” the researchers write. The first stage uses the text to create a “gist” of the video, basically a blurry image of the background with a blurry blob where the main action takes place. The second stage takes both the gist and the text and produces a short video. During training, a second network acts as a “discriminator.” It sees the video generated to illustrate, say, “sailing on the sea,” alongside a real video of sailing on the sea, and it is trained to pick the real one.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at sciencemag.org →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.