Home / Models / Realtime TTS-2

Realtime TTS-2

Maker: Inworld AI · speech synthesis

Inworld's flagship conversational voice model: natural-language voice direction, conditioning on prior dialogue audio, one voice identity across 100+ languages (docs claim 200+ languages and locales), ~100 ms TTFB.

Understands

text, audio

Produces

speech

Contents verified with the vendor: 2026-08-18

Included in subscriptions

No catalogue entry marks this model as part of a plan yet. It will appear here once the plan contents are confirmed with the vendor.