🤖 AI & Beyond

A revolutionary AI translation system for headphones replicates multiple voices at once.

Here’s the rewritten content with the specified adjustments:

Spatial Speech Translation consists of two AI models. The first model divides the space surrounding the person wearing the headphones into small regions and uses a neural network to search for potential speakers and pinpoint their direction.

The second model translates the speakers’ words from French, German, or Spanish into English text using publicly available data sets. This model also extracts the unique characteristics and emotional tone of each speaker’s voice, such as pitch and amplitude, and applies those properties to the text, effectively creating a “cloned” voice. Consequently, when the translated version of a speaker’s words is relayed to the headphone wearer a few seconds later, it sounds as if it’s coming from the speaker’s direction, with a voice that closely resembles the original, rather than a robotic tone.

Given the challenges of isolating human voices, incorporating this capability into a real-time translation system, accurately mapping the distance between the wearer and the speaker, and achieving minimal latency on an actual device is noteworthy. Samuele Cornell, a postdoctoral researcher at Carnegie Mellon University’s Language Technologies Institute, who did not work on the project, emphasizes that “real-time speech-to-speech translation is incredibly hard.” He notes, “Their results are very good in the limited testing settings. However, for a real product, much more training data would be needed—potentially including noise and real-world recordings from the headset, rather than purely relying on synthetic data.”

Gollakota’s team is currently working on reducing the time it takes for the AI translation to activate after a speaker says something, which will support more natural-sounding conversations between individuals speaking different languages. “We want to significantly decrease that latency to less than a second, so you can maintain the conversational vibe,” Gollakota explains.

This remains a considerable challenge, as the speed at which an AI system can translate one language into another is influenced by the structural differences of the languages. Among the three languages the Spatial Speech Translation was trained on, the system responds most quickly when translating French into English, followed by Spanish, and then German—reflecting how German, unlike the others, places a sentence’s verbs and much of its meaning at the end rather than at the beginning, according to Claudio Fantinuoli, a researcher at Johannes Gutenberg University of Mainz in Germany, who did not work on the project.

Reducing the latency could impact the accuracy of translations, he cautions: “The longer you wait [before translating], the more context you have, resulting in better translation. It’s a balancing act.”

Let me know if you need any further adjustments!