Voice turns text into spoken audio — narration, a voiceover line, or a line of dialogue — separate from Music Generation, which writes and performs a song. New voice generations use xAI; existing Qwen generations remain restorable as legacy history.
For pricing details, see the Pricing page.
Providers
| Provider | Description |
|---|---|
| xAI | Active provider for expressive multilingual speech with 26 built-in voices and inline delivery tags |
| Qwen TTS (Alibaba) | Deprecated legacy provider, retained only for restoring existing generations |
How it works
Select Voice
Open Create and select Voice. xAI is selected for new generations.
Write the line
Type the text you want spoken and configure the voice, language, speed, normalization, and output quality.
Generate
Generate the voiceover. MP3 profiles expose a Live preview control on the pending card as soon as audio starts arriving, while the same job continues saving the final result to your gallery. Studio WAV uses the completed gallery player.
xAI controls
- 26 built-in voices
- Auto-detection or 20 explicit languages
- Speech speed from 0.7× to 1.5×
- Optional normalization of numbers, abbreviations, and symbols into spoken form
- Standard MP3, high-fidelity MP3, and studio WAV output profiles
- Expressive tags such as
[pause],[laugh], or<whisper>quiet words</whisper>directly in the text - Up to 15,000 characters per generation
Tips
- Keep punctuation clean — commas and periods control pacing the same way they do when read aloud
- Use Voice for narration or dialogue lines; use Music when you want a full musical performance instead of spoken audio