Meta Voicebox Unveiled as New Text-to-Speech Generative AI Model: All Details

Meta's Voicebox is claimed to deliver audio clips using a two-second audio sample.

Advertisement
Written by Nithya P Nair, Edited by Siddharth Suvarna | Updated: 20 June 2023 14:30 IST
Highlights
  • Meta Platforms introduced Voicebox
  • Voicebox is a machine-learning model that can generate speech
  • It supports six languages
Meta Voicebox Unveiled as New Text-to-Speech Generative AI Model: All Details

Voicebox claimed to generate audio samples 20 times faster than Microsoft's VALL-E

Photo Credit: Meta

Meta announced Voicebox, its advanced artificial intelligence (AI) tool that can generate speech from text last week. The latest tool by Facebook parent Meta is claimed to produce high-quality audio clips and edit pre-recorded audio while preserving the content and style of the audio. It is said to be multilingual and claimed to deliver speech in six languages. The machine learning model can be used for noise removal as well. Meta's Voicebox also has the ability to replace misspoken words without having to re-record an entire speech. The new generative text-to-speech model works like the new AI innovations including ChatGPT and Dall-E.

Facebook's parent company Meta unveiled Voicebox via a blog post last week. This new generative AI model can perform speech generation tasks — like editing, sampling, and stylising. It is claimed to deliver audio clips from a two-second audio sample and edit pre-recorded audio while keeping the content and style of the audio.

The text-to-speech model is promised to perform tasks like noise removal, content editing, style conversion, and diverse sample generation. It is stated to modify any part of a given sample and recreate a portion of the speech that's interrupted by noise such as car horns or barking dogs. The AI model can also be used to replace misspoken words without having to re-record an entire speech.

Voicebox can synthesise speech across six languages — English, French, Spanish, German, Polish, and Portuguese. It can create a reading of the text in any of those languages, even when the sample speech and the text are in different languages.

Advertisement

Voicebox claimed to outperform Microsoft's VALL-E and generate audio samples 20 times faster. "Our results show that speech recognition models trained on Voicebox-generated synthetic speech perform almost as well as models trained on real speech, with 1 percent error rate degradation as opposed to 45 to 70 percent degradation with synthetic speech from previous text-to-speech models", Meta AI detailed in a research paper. Further, a few audio samples are listed to show users the working of Voicebox.

In the blog, Meta further claims that Voicebox can generate speech that is more representative of how people talk in the real world in the aforementioned six languages. The company believes that this capability could be used to generate synthetic data to help better train a speech assistant model in the near future.

Advertisement

Voicebox is currently under development and is not available to public users. Meta says it realises that this technology brings the potential for misuse and unintended harm like the current AI innovations. It is said to be working on an effective classifier that can distinguish between authentic speech and audio generated with Voicebox to mitigate these possible future risks.


Apple unveiled its first mixed reality headset, the Apple Vision Pro, at its annual developer conference, along with new Mac models and upcoming software updates. We discuss all the most important announcements made by the company at WWDC 2023 on Orbital, the Gadgets 360 podcast. Orbital is available on Spotify, Gaana, JioSaavn, Google Podcasts, Apple Podcasts, Amazon Music and wherever you get your podcasts.
Affiliate links may be automatically generated - see our ethics statement for details.
 

For the latest tech news and reviews, follow Gadgets 360 on X, Facebook, WhatsApp, Threads and Google News. For the latest videos on gadgets and tech, subscribe to our YouTube channel. If you want to know everything about top influencers, follow our in-house Who'sThat360 on Instagram and YouTube.

Advertisement

Related Stories

Popular Mobile Brands
  1. Vivo Y400 Pro 5G With 5,500mAh Battery Launched in India: Price, Features
  2. Samsung Galaxy M36 5G India Launch Date and Key Features Revealed
  3. Nothing Phone 3 to Get New Glyph Matrix Interface on the Rear Panel
  4. OTT Releases This Week: Ground Zero, Detective Sherdil, Found S2, and More
  5. Samsung Galaxy S25 FE Leaked Render Suggests Improved Design
  6. Adobe Launches a New Camera App for iPhone With Full Manual Controls
  7. Nothing Headphone 1 Renders Leaked Ahead of July 1 Launch: See Design
  8. Oppo Find X9 Pro Leak Suggests Potential Camera Specifications
  9. Vivo Y400 Pro 5G India Launch Today: All You Need to Know
  10. 16 Billion Login Credentials Have Been Leaked in Massive Data Breach
  1. Gigabyte Aorus Master 16 AI PC With Intel Core Ultra 9 Chip, Up to GeForce RTX 5080 GPU Launched in India
  2. Vivo X Fold 5 India Launch Reportedly Set for Mid-July
  3. Trump Extends Deadline for US TikTok Sale to September
  4. Nothing Headphone 1 Renders and Live Images Leak Ahead of July 1 Launch; Shows Unique Design
  5. BBC Said to Have Threatened Legal Action Against AI Start-up Perplexity Over Content Scraping
  6. Adobe Launches Project Indigo, a Camera App for iPhone With Full Manual Controls
  7. Oppo Find X9 Pro Camera Details Leaked; Said to Feature Samsung ISOCELL HP5 Sensor
  8. Nintendo Switch 2 Third-Party Game Sales Reportedly 'Very Low' Despite Console's Record Launch
  9. 16 Billion Login Credentials Leaked in Massive Data Breach Impacting Apple, Google and More
  10. Vivo Y400 Pro 5G With 50-Megapixel Rear Camera, 5,500mAh Battery Launched in India: Price, Specifications
Gadgets 360 is available in
Download Our Apps
Available in Hindi
© Copyright Red Pixels Ventures Limited 2025. All rights reserved.