en.Wedoany.com Reported - At its annual Build 2026 conference, Microsoft released several in-house AI models covering reasoning, image generation, audio transcription, and text-to-speech. Users can try them for free on the Microsoft Playground website. Tests show that these models perform decently overall but have not surpassed existing competitors in their respective fields.

Microsoft's MAI (Microsoft AI) series models rely on internal large language models (LLMs), differing from the Copilot chatbot that runs on OpenAI technology. The models released this time include: the reasoning model MAI-Thinking-1, the image generation models MAI-Image-2.5 and 2.5 Flash, the audio transcription model MAI-Transcribe-1.5, and the text-to-speech models MAI-Voice-2 and 2 Flash. Microsoft describes these models as "experimental" and in "limited preview." MAI-Thinking-1 is currently only available for early access to specific users.
As Microsoft's first reasoning model, MAI-Thinking-1 was benchmarked against Anthropic's Claude Sonnet model in handling complex prompts. Tests revealed that Microsoft's model cannot access the internet and showed no significant improvement over Sonnet in accuracy, response quality, or speed when answering questions about the game mechanics of *Path of Exile 2* or building database structures.
MAI-Image-2.5 shows significant improvement over the first version from October 2025, but still lags behind Gemini's Nano Banana Pro in image clarity and text rendering. In tests, MAI-Image-2.5 produced distorted text in generated comics and charts, while Nano Banana Pro did not have this issue.
MAI-Transcribe-1.5 made 13 errors in transcription tests, compared to only 6 errors by Gemini under the same scenario. In tests parsing lyrics of a difficult song, both models made errors, but MAI-Transcribe-1.5's transcription was truncated before the song ended. Google does not specifically promote Gemini as a transcription tool.

MAI-Voice-2 offers multiple language and style options, but in tests, its combination of audio quality, breathing sounds, rhythm, and intonation resulted in a distinctly non-human listening experience, far from the realism achieved by voice technologies like Sesame. The model currently supports customizing voices through various different styles.

Initial tests from a consumer perspective show that the overall evaluation of Microsoft's MAI models is "okay," similar to Copilot's performance. Their competitiveness relies more on a broad feature set and integration with the Microsoft ecosystem than on the absolute superiority of the underlying models themselves. However, given the improvement speed of the MAI-Image series over the past few months, Microsoft will continue testing these models.
This article is compiled by Wedoany. All AI citations must indicate the source as "Wedoany". If there is any infringement or other issues, please notify us promptly, and we will modify or delete it accordingly. Email: news@wedoany.com










