MAI-Voice-2 is an expressive text-to-speech model from Microsoft AI. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18 locales, fine-grained control of tone and delivery, multi-speaker generation, and voice prompting from short audio clips without fine-tuning. The model prioritizes naturalness and expressivity over latency-critical generation.
Modalities
Price
$22/M characters
Released
Jun 2, 2026
This model is hosted by one provider. OpenRouter forwards every request to it directly — no routing decisions to make.
Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better).
Uptime is the percentage of the past 3 days that at least one provider was responding to requests. Availability is the percentage of time that inference was successfully served. OpenRouter continuously monitors and uses the next-best provider when one returns an error.
Scores on standardized evaluations. Higher percentages are better — and rank percentile shows where this model lands among all models on OpenRouter.
| Source | Benchmark | Score |
|---|---|---|
| Design Arena | MAI-Voice-2 Models Arena Audiorealism Elo | 1228 |
| Design Arena | MAI-Voice-2 Models Arena Text-To-Speech Elo | 1191 |
Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.
Token volume and request traffic to this model over time.
Drop-in code to call this model. OpenRouter's API is OpenAI-compatible — most SDKs work by just swapping the base URL. The only thing that changes between models is the model slug below.
| $22.00 | 1.15s |
Latency
1.15s
P50, best provider
100.00%
97.18%
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.
MAI-Voice-2 is an expressive text-to-speech model from Microsoft AI. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18 locales, fine-grained control of tone and delivery, multi-speaker generation, and voice prompting from short audio clips without fine-tuning.
MAI-Voice-2 costs $22.00/M characters.
MAI-Voice-2 ships 4 voices. Pass a voice ID in the voice field of a text-to-speech request; the IDs this endpoint accepts are listed with the model in the models API.
MAI-Voice-2 accepts text as input and returns generated speech audio.
MAI-Voice-2-Flash is another speech model from Microsoft.
MAI-Voice-2 was released on June 2, 2026.