Home / Models / MAI-Voice-2

MAI-Voice-2

Maker: Microsoft · speech synthesis

Microsoft AI's highest-fidelity text-to-speech: 15 languages and 18 locales, prebuilt voices, emotion control via SSML, long-form output with a consistent speaker. Voice cloning from a 5–60 second sample is granted on application.

доступна как публичное превью в Azure Speech; клонирование голоса — по отдельной заявке

Understands

text, audio

Produces

speech

Contents verified with the vendor: 2026-08-18

Included in subscriptions

No catalogue entry marks this model as part of a plan yet. It will appear here once the plan contents are confirmed with the vendor.