Microsoft previews MAI-Image-2.5-Pro and MAI-Voice-2-Flash

Microsoft has introduced two new models under its MAI family through Azure AI Foundry: MAI-Image-2.5-Pro and MAI-Voice-2-Flash. The former targets high-fidelity visual output, while the latter is positioned as a faster, lower-cost option for speech generation. Both are in preview, meaning developers with Foundry access can begin evaluating them before any general release.
MAI-Image-2.5-Pro sits in Microsoft's image generation lineup as a model built for quality-focused use cases - scenarios where detail and visual accuracy take priority over speed or cost. The "2.5-Pro" naming follows a convention similar to what other labs use to distinguish capable, higher-tier models from their more lightweight counterparts, suggesting this is intended as a more capable offering within the MAI image family.
MAI-Voice-2-Flash, on the other hand, carries the "Flash" label that has become shorthand across the industry for models optimized for latency and cost rather than maximum quality. For developers building applications where speech needs to be generated quickly and at scale - such as voice assistants, automated call systems, or real-time narration - a faster, cheaper speech model can meaningfully affect the practicality of deployment.
The releases fit into a broader pattern of Microsoft developing its own first-party models to complement the third-party offerings already available through Azure AI Foundry and its partnership with OpenAI. Having proprietary models gives Microsoft more control over pricing, latency, and the ability to offer differentiated options to enterprise customers. With both models currently in preview, further details on benchmarks, pricing tiers, and availability timelines are likely to follow as testing progresses.

