Goto

Collaborating Authors

 speech2face


Artificial Intelligence Generates Humans' Faces Based on Their Voices

#artificialintelligence

A new neural network developed by researchers from the Massachusetts Institute of Technology is capable of constructing a rough approximation of an individual's face based solely on a snippet of their speech, a paper published in pre-print server arXiv reports. The team trained the artificial intelligence tool--a machine learning algorithm programmed to "think" much like the human brain--with the help of millions of online clips capturing more than 100,000 different speakers. Dubbed Speech2Face, the neural network used this dataset to determine links between vocal cues and specific facial features; as the scientists write in the study, age, gender, the shape of one's mouth, lip size, bone structure, language, accent, speed and pronunciation all factor into the mechanics of speech. According to Gizmodo's Melanie Ehrenkranz, Speech2Face draws on associations between appearance and speech to generate photorealistic renderings of front-facing individuals with neutral expressions. Although these images are too generic to identify as a specific person, the majority of them accurately pinpoint speakers' gender, race and age.


With this AI, your voice could give away your face

#artificialintelligence

In research published on Arxiv, a publishing site for non-peer-reviewed papers, MIT researchers created a way to reconstruct some people's very rough likeness based on a short audio clip. The paper, "Speech2Face: Learning the Face Behind a Voice," explains how they took a dataset made up of millions of clips from YouTube and created a neural network-based model that learns vocal attributes associated with facial features from the videos. Now, when the system hears a new sound bite, the AI can use what it's learned to guess what the face might look like. The researchers, led by MIT postdoctoral student Tae-Hyun Oh, do briefly acknowledge the privacy concerns in the paper, explaining in an "Ethical Consideration" section that Speech2Face was trained to capture visual features like gender and age that are common, and only when there was enough evidence from the voice to do so. In other words, the system is not trying or able to produce images of specific people.