Skip to main content
POST /v1/audio serves both speech and music; the model decides which, and the fields you send differ.

Speech

Send text and a voice_name (or a voice id from your workspace). ElevenLabs Multilingual v2 and Turbo v2.5 cover most languages; v3 understands expressive cues in the text. Long scripts can be split into segments with different voices for dialogue.

Music

Send a prompt describing style, mood and instrumentation. Lyria takes tag-style prompts with energy and tempo; ElevenLabs Music (music_v1 / music_v2) takes free text, an instrumental toggle, and lyrics with [Verse] / [Chorus] markers. Pricing is per character for speech and per generation for music; estimate with character_count for speech. Voice-specific tips: the ElevenLabs prompting guide.