Sound & Voice Parameters
Generational step-up multimodal text-to-speech from Google Gemini
Multilingual zero-shot voice synthesis from Mistral AI
Lightweight open-weight model supporting 8 languages & 50+ speakers
Expressive speech model from xAI across 20+ languages
Highly expressive Microsoft Azure voices with style controls
Low latency 24kHz mono voice model for interactive real-time apps
Ultra low-latency natural human speech across 7 languages
Alibaba Qwen fast high-fidelity speech synthesis model
Alibaba Qwen enhanced speech synthesis model
Pure transformer TTS model from Zyphra
Hybrid architecture TTS model from Zyphra
Fine-tuned for natural prosody and expressive English delivery
Conversational speech model for natural dialogue & read-speech
MiniMax high-definition voice generator
MiniMax ultra-fast voice generator