Skip to main content
POST
Generate speech or music

Authorizations

Authorization
string
header
required

An organization API key created in Settings → Team, sent as Authorization: Bearer <key>.

Headers

Idempotency-Key
string | null

Reuse the same value when retrying the same request; a repeat returns the original generation without charging again. Reusing it with a different body returns 409.

Maximum string length: 128

Body

application/json

Generate speech or music; the model decides which.

apply_text_normalization
string | null

Text normalization: on/off/auto

bpm
integer | null

Optional BPM override (60-200). If not provided, tempo is used to estimate. (music generation only)

Required range: 60 <= x <= 200
composition_plan
Composition Plan · object | null

ElevenLabs MusicPrompt composition plan. Mutually exclusive with prompt/prompt_tags; section durations control track length.

content_type
string | null

Content type preset name

duration
integer
default:30

Duration in seconds. Lyria allows 5-30 (validated per-model in the route); ElevenLabs Music allows up to 300. Not used for TTS.

Required range: 3 <= x <= 600
energy
string
default:medium

Energy level: 'low', 'medium', or 'high' (music generation only)

force_instrumental
boolean | null
default:false

Generate instrumental-only music (ElevenLabs Music, prompt mode only)

idempotency_key
string | null

Client-supplied key (≤128 chars); a repeat with the same key returns the original generation without re-charging

Maximum string length: 128
instrumental
boolean | null

Generate instrumental-only music with no vocals (Lyria 3 only). Mutually exclusive with lyrics.

language
string | null

Language for TTS synthesis: Qwen3 uses names such as 'Auto', 'English', 'Chinese'; Gemini uses BCP-47 codes. ElevenLabs auto-detects language and ignores this shared field.

lyrics
string | null

Custom lyrics for vocal generation (Lyria 3 only). Use [Verse], [Chorus], [Bridge] markers. Pro model also supports [MM:SS] timestamps.

model
string | null

Audio model identifier. Music: 'models/lyria-realtime-exp'. TTS Gemini: 'gemini-3.1-flash-tts-preview'. TTS ElevenLabs: 'eleven_v4', 'eleven_v3', 'eleven_multilingual_v2', 'eleven_turbo_v2_5'

negative_prompt
string | null

What to avoid in the music (Lyria 2 only). Ignored by Lyria RealTime.

next_audio_gen_id
string | null

Generation ID for next audio context

next_context_mode
string | null
default:text

'text' or 'audio'

next_request_ids
string[] | null

ElevenLabs request IDs for next audio

next_text
string | null

Next text context (not spoken)

output_format
string | null

ElevenLabs MP3 or WAV output format (e.g. mp3_44100_128, wav_24000)

previous_audio_gen_id
string | null

Generation ID for previous audio context

previous_context_mode
string | null
default:text

'text' or 'audio'

previous_request_ids
string[] | null

ElevenLabs request IDs for previous audio

previous_text
string | null

Previous text context (not spoken)

project_id
string | null

Project scope for the generated asset (optional)

prompt
string | null

Free-text music prompt (ElevenLabs Music). Mutually exclusive with composition_plan.

prompt_tags
string[] | null

List of prompt tags describing the desired music (e.g. ['energetic', 'rock', 'guitar']). Required for music generation.

pronunciation_dictionary_ids
string[] | null

Dictionary entry IDs to apply

qwen3_embedding_id
string | null

Qwen3 voice embedding ID (from voice profile clone)

qwen3_model_size
string | null

Qwen3 model size

sample_count
integer | null

Number of variations per API call (Lyria 2 only, 1-4). Mutually exclusive with seed.

Required range: 1 <= x <= 4
seed
integer | null

Random seed for reproducible output (Lyria 2 only). Mutually exclusive with sample_count > 1.

segments
SpeechSegment · object[] | null

Multi-speaker script as an ordered list of {voice_name, text} lines. Mutually exclusive with top-level text. ElevenLabs v3 without stitching context and two-voice Gemini conversations use native dialogue; other requests synthesize turns independently. All return one audio file.

similarity_boost
number | null

Voice clarity/similarity for ElevenLabs (0.0-1.0). Higher = closer to original voice.

Required range: 0 <= x <= 1
speed
number | null

Speech speed. Per-model range comes from the catalog.

Required range: 0.7 <= x <= 1.3
stability
number | null

Voice stability for ElevenLabs (0.0-1.0). Higher = more consistent, lower = more expressive.

Required range: 0 <= x <= 1
style
number | null

Style exaggeration (0.0-1.0)

Required range: 0 <= x <= 1
temperature
number | null

Optional creativity/randomness (0.0-2.0). Higher = more variation. (music generation only)

Required range: 0 <= x <= 2
tempo
string
default:medium

Tempo: 'slow', 'medium', or 'fast' (music generation only)

text
string | null

Text to convert to speech. Required for single-speaker TTS. Mutually exclusive with segments.

voice_display_name
string | null

Human-readable voice name for UI display (voice_name carries the voice_id for backend)

voice_name
string | null

Voice name for TTS. Gemini: 'Zephyr', 'Charon', 'Kore', 'Fenrir', 'Aoede', 'Puck'. ElevenLabs: a current premade name such as 'Sarah', 'George' or 'Laura', or any library voice id.

voice_profile_id
string | null

Voice profile ID (used for Qwen3 cloned voices)

Response

Successful Response

One generation. Poll GET /v1/generations/{id} until status is completed or failed.

created_at
string<date-time>
required
id
string
required
Example:

"6f1c2d3e-9a0b-4c7d-8e1f-2a3b4c5d6e7f"

model
string
required
Example:

"seedance-2-5-fal"

prompt
string
required
status
enum<string>
required
Available options:
pending,
completed,
failed
type
enum<string>
required
Available options:
image,
video,
audio
completed_at
string<date-time> | null
credits_charged
number | null

Credits actually charged; null while pending.

error
GenerationError · object | null
idempotency_key
string | null
outputs
GenerationOutput · object[]