Realtime TTS-2
Our flagship, top-ranked model — the best choice for production
- Best quality and steerability, with natural language steering for more contextually aware speech
- Support for 200+ languages and locales
- Ultra-low latency (~120ms median latency) at high concurrency
- High quality instant voice cloning
- Enhanced timestamps with phonetic details and visemes
Realtime TTS 1.5 Max
Rich, expressive speech with maximum stability
- Support for 15 languages
- Optimized for real-time use (<200ms median latency)
- High quality instant voice cloning
Realtime TTS 1.5 Mini
Our most cost-efficient model — for English workloads where price is the top priority
- Lowest cost per character
- Best suited to English; use TTS-2 for other languages
- High quality instant voice cloning
Models overview
Looking for
inworld-tts-1 or inworld-tts-1-max? These previous-generation models were discontinued on June 15, 2026. Requests to them are now automatically routed to their 1.5 successors (inworld-tts-1 → inworld-tts-1.5-mini, inworld-tts-1-max → inworld-tts-1.5-max). We recommend migrating to inworld-tts-2 to improve quality and latency.