Microsoft’s New AI Models Go Beyond Just Text

Microsoft is doubling down on AI models that aren’t large language models. The company announced on Thursday that it’s releasing three new models: brand new models for voice and text transcription, and the second generation of its in-house image model.

The voice and text transcription models are the first of their kind from Microsoft. The transcription model can translate recordings into text in 25 different languages. It’s built for video captioning, meeting transcription and voice agents. The voice model can create audio recordings up to 60 seconds long. The company says its second-generation image model has a faster generation speed and more lifelike depictions, improving on its previous model. They’re available now in Microsoft’s Foundry and MAI playground, with future plans to bring MAI-Image-2 to Bing and PowerPoint. Developers can check out pricing info here.

These new models are a clear sign that Microsoft is looking to expand its offerings across the AI

...

Keep reading this article on CNET.