Microsoft’s latest audio generation in Copilot is powered by the homegrown MAI-Voice-1 AI model.
Photo Credit: Microsoft
The new Copilot feature is available to all users of the platform via a personal account
Microsoft is adding another new artificial intelligence (AI) feature to Copilot, giving it the ability to natively generate audio. On Wednesday, the Redmond-based tech giant announced that Copilot is getting a new audio generation feature where users will be able to hand it a script and it will convert it into an AI voiceover in different styles. Since it is native voice generation, none of the modes will sound like typical text-to-speech models. Notably, the company is powering this capability via the homegrown MAI-Voice-1 AI model.
In a post on X (formerly known as Twitter), Mustafa Suleyman, CEO of Microsoft AI, announced the release of Copilot's new audio generation modes. He highlighted that these are powered by the MAI-Voice-1 AI model, which was released at the end of August. Currently, this experience is only available via Copilot Labs when signing in using a personal account.
You asked, we shipped! Scripted mode just dropped for audio generation in Copilot Labs (c/o our new MAI-Voice-1 model).
— Mustafa Suleyman (@mustafasuleyman) September 10, 2025
Scripted mode: reads your input verbatim
Emotive: riffs a bit for max drama
Story: performs multiple voices/characters
Try out all 3 ➡️ https://t.co/9hL81LTFwF pic.twitter.com/rOVZKGbDjX
There are three modes to try out. First is the Scripted mode, where the AI chatbot reads out the input verbatim, without adding any unnecessary flair or style. These are best used for tasks such as formal announcements, document narration, and information presentation.
The second mode is dubbed Emotive. Suleyman says it is more focused on making the input sound dramatic and flashy. The voice here will include a wide range of intonation, pitch, and tone to deliver a performative piece. This is ideal for advertising, marketing, or informal narration.
Copilot's final audio generation mode is Story. This is the most versatile format, which includes multiple voices and characters. The company says this mode is ideal for storytelling, podcast-like presentations, and analysis-related tasks. The feature is currently free to use, although Microsoft has not mentioned any rate limits. It is unclear when the feature will be released into the Copilot mobile and desktop apps.
Notably, at the time of release, Microsoft said the MAI-Voice-1 is a speech generation model that natively generates expressive and natural-sounding voice. It can generate a full minute of audio in under a second on a single GPU. The tech giant trained the model on around 15,000 Nvidia GPUs.
Get your daily dose of tech news, reviews, and insights, in under 80 characters on Gadgets 360 Turbo. Connect with fellow tech lovers on our Forum. Follow us on X, Facebook, WhatsApp, Threads and Google News for instant updates. Catch all the action on our YouTube channel.
The Rookie Season 7 OTT Release Date: When and Where to Watch it Online?
Dominic and the Ladies' Purse OTT Release Date: When and Where to Watch it Online?
Kesariya at 100 Season 1 Now Streaming on ZEE5: When and Where to Watch Docuseries Online?
Radhika Apte’s New Psychological Thriller Saali Mohabbat Now Streaming on ZEE5