> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ekly.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Generate speech or music

> Submits a speech or music generation; the model decides which. Returns immediately with `status: pending`; credits are reserved now and settled when the job finishes. Poll `GET /v1/generations/{id}` until `completed` or `failed`.



## OpenAPI

````yaml /openapi.v1.json post /v1/audio
openapi: 3.1.0
info:
  contact:
    email: hello@ekly.ai
    name: Ekly
    url: https://docs.ekly.ai
  description: >-
    Generate images, video, music and speech with the same models and credits as
    the Ekly app.


    Authenticate with an organization API key from Settings → Team as a bearer
    token. Every endpoint lives under /v1 and follows the additive-only
    versioning policy at https://docs.ekly.ai/versioning.
  termsOfService: https://ekly.ai/terms
  title: Ekly API
  version: '1.0'
servers:
  - description: Production
    url: https://api.ekly.ai
security:
  - ApiKey: []
tags:
  - description: Public, versioned developer API
    name: v1
paths:
  /v1/audio:
    post:
      tags:
        - v1
      summary: Generate speech or music
      description: >-
        Submits a speech or music generation; the model decides which. Returns
        immediately with `status: pending`; credits are reserved now and settled
        when the job finishes. Poll `GET /v1/generations/{id}` until `completed`
        or `failed`.
      operationId: create_audio
      parameters:
        - description: >-
            Reuse the same value when retrying the same request; a repeat
            returns the original generation without charging again. Reusing it
            with a different body returns 409.
          in: header
          name: Idempotency-Key
          required: false
          schema:
            anyOf:
              - maxLength: 128
                type: string
              - type: 'null'
            description: >-
              Reuse the same value when retrying the same request; a repeat
              returns the original generation without charging again. Reusing it
              with a different body returns 409.
            title: Idempotency-Key
      requestBody:
        content:
          application/json:
            example:
              model: eleven_multilingual_v2
              text: Welcome back. Your export is ready.
              voice_name: Rachel
            schema:
              $ref: '#/components/schemas/CreateAudioRequest'
        required: true
      responses:
        '200':
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/GenerationResponse'
          description: Successful Response
        '401':
          content:
            application/json:
              example:
                error:
                  code: unauthenticated
                  message: >-
                    This API key is not valid. Create a new one in Settings →
                    Team.
                  trace_id: c3746d5302dc4be6b2e8fa99fa762a8c
              schema:
                $ref: '#/components/schemas/ErrorResponse'
          description: API key missing, invalid, revoked or expired.
        '402':
          content:
            application/json:
              example:
                error:
                  code: payment_required
                  message: >-
                    Insufficient credits: this generation needs 39.6 credits and
                    12.0 are available.
                  trace_id: c3746d5302dc4be6b2e8fa99fa762a8c
              schema:
                $ref: '#/components/schemas/ErrorResponse'
          description: Not enough credits.
        '403':
          content:
            application/json:
              example:
                error:
                  code: forbidden
                  message: API keys can only call the /v1 API.
                  trace_id: c3746d5302dc4be6b2e8fa99fa762a8c
              schema:
                $ref: '#/components/schemas/ErrorResponse'
          description: Not allowed for this key or organization.
        '409':
          content:
            application/json:
              example:
                error:
                  code: conflict
                  message: >-
                    This idempotency key was already used with a different
                    request.
                  trace_id: c3746d5302dc4be6b2e8fa99fa762a8c
              schema:
                $ref: '#/components/schemas/ErrorResponse'
          description: Idempotency key reused with a different request.
        '422':
          content:
            application/json:
              example:
                error:
                  code: validation_error
                  details:
                    - loc:
                        - body
                        - model
                      msg: Field required
                      type: missing
                  message: Request validation failed.
                  trace_id: c3746d5302dc4be6b2e8fa99fa762a8c
              schema:
                $ref: '#/components/schemas/ErrorResponse'
          description: Request failed validation.
        '429':
          content:
            application/json:
              example:
                error:
                  code: rate_limited
                  message: Rate limit exceeded; retry after 12s.
                  trace_id: c3746d5302dc4be6b2e8fa99fa762a8c
              schema:
                $ref: '#/components/schemas/ErrorResponse'
          description: Rate limit exceeded; see Retry-After.
      security:
        - ApiKey: []
components:
  schemas:
    CreateAudioRequest:
      additionalProperties: false
      description: Generate speech or music; the model decides which.
      properties:
        apply_text_normalization:
          anyOf:
            - type: string
            - type: 'null'
          description: 'Text normalization: on/off/auto'
          title: Apply Text Normalization
        bpm:
          anyOf:
            - maximum: 200
              minimum: 60
              type: integer
            - type: 'null'
          description: >-
            Optional BPM override (60-200). If not provided, tempo is used to
            estimate. (music generation only)
          title: Bpm
        composition_plan:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          description: >-
            ElevenLabs MusicPrompt composition plan. Mutually exclusive with
            prompt/prompt_tags; section durations control track length.
          title: Composition Plan
        content_type:
          anyOf:
            - type: string
            - type: 'null'
          description: Content type preset name
          title: Content Type
        duration:
          default: 30
          description: >-
            Duration in seconds. Lyria allows 5-30 (validated per-model in the
            route); ElevenLabs Music allows up to 300. Not used for TTS.
          maximum: 600
          minimum: 3
          title: Duration
          type: integer
        energy:
          default: medium
          description: 'Energy level: ''low'', ''medium'', or ''high'' (music generation only)'
          title: Energy
          type: string
        force_instrumental:
          anyOf:
            - type: boolean
            - type: 'null'
          default: false
          description: >-
            Generate instrumental-only music (ElevenLabs Music, prompt mode
            only)
          title: Force Instrumental
        idempotency_key:
          anyOf:
            - maxLength: 128
              type: string
            - type: 'null'
          description: >-
            Client-supplied key (≤128 chars); a repeat with the same key returns
            the original generation without re-charging
          title: Idempotency Key
        instrumental:
          anyOf:
            - type: boolean
            - type: 'null'
          description: >-
            Generate instrumental-only music with no vocals (Lyria 3 only).
            Mutually exclusive with lyrics.
          title: Instrumental
        language:
          anyOf:
            - type: string
            - type: 'null'
          description: >-
            Language for TTS synthesis: Qwen3 uses names such as 'Auto',
            'English', 'Chinese'; Gemini uses BCP-47 codes. ElevenLabs
            auto-detects language and ignores this shared field.
          title: Language
        lyrics:
          anyOf:
            - type: string
            - type: 'null'
          description: >-
            Custom lyrics for vocal generation (Lyria 3 only). Use [Verse],
            [Chorus], [Bridge] markers. Pro model also supports [MM:SS]
            timestamps.
          title: Lyrics
        model:
          anyOf:
            - type: string
            - type: 'null'
          description: >-
            Audio model identifier. Music: 'models/lyria-realtime-exp'. TTS
            Gemini: 'gemini-3.1-flash-tts-preview'. TTS ElevenLabs: 'eleven_v4',
            'eleven_v3', 'eleven_multilingual_v2', 'eleven_turbo_v2_5'
          title: Model
        negative_prompt:
          anyOf:
            - type: string
            - type: 'null'
          description: >-
            What to avoid in the music (Lyria 2 only). Ignored by Lyria
            RealTime.
          title: Negative Prompt
        next_audio_gen_id:
          anyOf:
            - type: string
            - type: 'null'
          description: Generation ID for next audio context
          title: Next Audio Gen Id
        next_context_mode:
          anyOf:
            - type: string
            - type: 'null'
          default: text
          description: '''text'' or ''audio'''
          title: Next Context Mode
        next_request_ids:
          anyOf:
            - items:
                type: string
              type: array
            - type: 'null'
          description: ElevenLabs request IDs for next audio
          title: Next Request Ids
        next_text:
          anyOf:
            - type: string
            - type: 'null'
          description: Next text context (not spoken)
          title: Next Text
        output_format:
          anyOf:
            - type: string
            - type: 'null'
          description: ElevenLabs MP3 or WAV output format (e.g. mp3_44100_128, wav_24000)
          title: Output Format
        previous_audio_gen_id:
          anyOf:
            - type: string
            - type: 'null'
          description: Generation ID for previous audio context
          title: Previous Audio Gen Id
        previous_context_mode:
          anyOf:
            - type: string
            - type: 'null'
          default: text
          description: '''text'' or ''audio'''
          title: Previous Context Mode
        previous_request_ids:
          anyOf:
            - items:
                type: string
              type: array
            - type: 'null'
          description: ElevenLabs request IDs for previous audio
          title: Previous Request Ids
        previous_text:
          anyOf:
            - type: string
            - type: 'null'
          description: Previous text context (not spoken)
          title: Previous Text
        project_id:
          anyOf:
            - type: string
            - type: 'null'
          description: Project scope for the generated asset (optional)
          title: Project Id
        prompt:
          anyOf:
            - type: string
            - type: 'null'
          description: >-
            Free-text music prompt (ElevenLabs Music). Mutually exclusive with
            composition_plan.
          title: Prompt
        prompt_tags:
          anyOf:
            - items:
                type: string
              type: array
            - type: 'null'
          description: >-
            List of prompt tags describing the desired music (e.g. ['energetic',
            'rock', 'guitar']). Required for music generation.
          title: Prompt Tags
        pronunciation_dictionary_ids:
          anyOf:
            - items:
                type: string
              type: array
            - type: 'null'
          description: Dictionary entry IDs to apply
          title: Pronunciation Dictionary Ids
        qwen3_embedding_id:
          anyOf:
            - type: string
            - type: 'null'
          description: Qwen3 voice embedding ID (from voice profile clone)
          title: Qwen3 Embedding Id
        qwen3_model_size:
          anyOf:
            - type: string
            - type: 'null'
          description: Qwen3 model size
          title: Qwen3 Model Size
        sample_count:
          anyOf:
            - maximum: 4
              minimum: 1
              type: integer
            - type: 'null'
          description: >-
            Number of variations per API call (Lyria 2 only, 1-4). Mutually
            exclusive with seed.
          title: Sample Count
        seed:
          anyOf:
            - type: integer
            - type: 'null'
          description: >-
            Random seed for reproducible output (Lyria 2 only). Mutually
            exclusive with sample_count > 1.
          title: Seed
        segments:
          anyOf:
            - items:
                $ref: '#/components/schemas/SpeechSegment'
              type: array
            - type: 'null'
          description: >-
            Multi-speaker script as an ordered list of {voice_name, text} lines.
            Mutually exclusive with top-level text. ElevenLabs v3 without
            stitching context and two-voice Gemini conversations use native
            dialogue; other requests synthesize turns independently. All return
            one audio file.
          title: Segments
        similarity_boost:
          anyOf:
            - maximum: 1
              minimum: 0
              type: number
            - type: 'null'
          description: >-
            Voice clarity/similarity for ElevenLabs (0.0-1.0). Higher = closer
            to original voice.
          title: Similarity Boost
        speed:
          anyOf:
            - maximum: 1.3
              minimum: 0.7
              type: number
            - type: 'null'
          description: Speech speed. Per-model range comes from the catalog.
          title: Speed
        stability:
          anyOf:
            - maximum: 1
              minimum: 0
              type: number
            - type: 'null'
          description: >-
            Voice stability for ElevenLabs (0.0-1.0). Higher = more consistent,
            lower = more expressive.
          title: Stability
        style:
          anyOf:
            - maximum: 1
              minimum: 0
              type: number
            - type: 'null'
          description: Style exaggeration (0.0-1.0)
          title: Style
        temperature:
          anyOf:
            - maximum: 2
              minimum: 0
              type: number
            - type: 'null'
          description: >-
            Optional creativity/randomness (0.0-2.0). Higher = more variation.
            (music generation only)
          title: Temperature
        tempo:
          default: medium
          description: 'Tempo: ''slow'', ''medium'', or ''fast'' (music generation only)'
          title: Tempo
          type: string
        text:
          anyOf:
            - type: string
            - type: 'null'
          description: >-
            Text to convert to speech. Required for single-speaker TTS. Mutually
            exclusive with segments.
          title: Text
        voice_display_name:
          anyOf:
            - type: string
            - type: 'null'
          description: >-
            Human-readable voice name for UI display (voice_name carries the
            voice_id for backend)
          title: Voice Display Name
        voice_name:
          anyOf:
            - type: string
            - type: 'null'
          description: >-
            Voice name for TTS. Gemini: 'Zephyr', 'Charon', 'Kore', 'Fenrir',
            'Aoede', 'Puck'. ElevenLabs: a current premade name such as 'Sarah',
            'George' or 'Laura', or any library voice id.
          title: Voice Name
        voice_profile_id:
          anyOf:
            - type: string
            - type: 'null'
          description: Voice profile ID (used for Qwen3 cloned voices)
          title: Voice Profile Id
      title: CreateAudioRequest
      type: object
    GenerationResponse:
      description: >-
        One generation. Poll GET /v1/generations/{id} until status is completed
        or failed.
      properties:
        completed_at:
          anyOf:
            - format: date-time
              type: string
            - type: 'null'
          title: Completed At
        created_at:
          format: date-time
          title: Created At
          type: string
        credits_charged:
          anyOf:
            - type: number
            - type: 'null'
          description: Credits actually charged; null while pending.
          title: Credits Charged
        error:
          anyOf:
            - $ref: '#/components/schemas/GenerationError'
            - type: 'null'
        id:
          examples:
            - 6f1c2d3e-9a0b-4c7d-8e1f-2a3b4c5d6e7f
          title: Id
          type: string
        idempotency_key:
          anyOf:
            - type: string
            - type: 'null'
          title: Idempotency Key
        model:
          examples:
            - seedance-2-5-fal
          title: Model
          type: string
        outputs:
          items:
            $ref: '#/components/schemas/GenerationOutput'
          title: Outputs
          type: array
        prompt:
          title: Prompt
          type: string
        status:
          enum:
            - pending
            - completed
            - failed
          title: Status
          type: string
        type:
          enum:
            - image
            - video
            - audio
          title: Type
          type: string
      required:
        - id
        - type
        - status
        - model
        - prompt
        - created_at
      title: GenerationResponse
      type: object
    ErrorResponse:
      description: Every non-2xx /v1 response.
      properties:
        error:
          $ref: '#/components/schemas/ErrorBody'
      required:
        - error
      title: ErrorResponse
      type: object
    SpeechSegment:
      description: >-
        One speaker line in a multi-speaker TTS script.


        Mutually exclusive with the top-level `text`/`voice_name` pair on

        AudioGenerateRequest — send either a single block or an ordered list of

        these, not both. ElevenLabs models with `supports.native_dialogue` use
        native

        dialogue unless request-stitching context is supplied. Gemini uses
        native dialogue for two distinct voices and per-turn

        synthesis for one distinct voice; other models synthesize independently.

        Every path returns one audio file / one AIGeneration row.
      properties:
        similarity_boost:
          anyOf:
            - maximum: 1
              minimum: 0
              type: number
            - type: 'null'
          title: Similarity Boost
        speed:
          anyOf:
            - maximum: 1.2
              minimum: 0.7
              type: number
            - type: 'null'
          title: Speed
        stability:
          anyOf:
            - maximum: 1
              minimum: 0
              type: number
            - type: 'null'
          title: Stability
        style:
          anyOf:
            - maximum: 1
              minimum: 0
              type: number
            - type: 'null'
          title: Style
        text:
          description: Spoken text for this line; the model's max_text_length applies
          maxLength: 20000
          minLength: 1
          title: Text
          type: string
        voice_display_name:
          anyOf:
            - type: string
            - type: 'null'
          description: Human-readable voice name for UI restore
          title: Voice Display Name
        voice_name:
          description: Provider voice id or name (same as top-level voice_name)
          minLength: 1
          title: Voice Name
          type: string
      required:
        - voice_name
        - text
      title: SpeechSegment
      type: object
    GenerationError:
      properties:
        code:
          description: >-
            One of the public failure categories: rate_limited, timeout, safety,
            invalid_voice, invalid_request, billing_or_access, provider_auth,
            provider_unavailable, provider_error.
          title: Code
          type: string
        message:
          title: Message
          type: string
      required:
        - code
        - message
      title: GenerationError
      type: object
    GenerationOutput:
      properties:
        asset_id:
          anyOf:
            - type: string
            - type: 'null'
          description: >-
            Workspace asset id for this output. Pass it as a reference input to
            chain generations.
          title: Asset Id
        content_type:
          anyOf:
            - type: string
            - type: 'null'
          examples:
            - video/mp4
          title: Content Type
        duration_seconds:
          anyOf:
            - type: number
            - type: 'null'
          title: Duration Seconds
        height:
          anyOf:
            - type: integer
            - type: 'null'
          title: Height
        url:
          description: >-
            Download URL. Fetch promptly or store the file; URLs are not
            permanent.
          title: Url
          type: string
        width:
          anyOf:
            - type: integer
            - type: 'null'
          title: Width
      required:
        - url
      title: GenerationOutput
      type: object
    ErrorBody:
      properties:
        code:
          description: >-
            Stable machine-readable code: unauthenticated, invalid_api_key,
            payment_required, forbidden, api_key_scope, not_found, conflict,
            validation_error, rate_limited, client_error, internal_error.
          title: Code
          type: string
        details:
          anyOf:
            - items:
                additionalProperties: true
                type: object
              type: array
            - type: 'null'
          description: >-
            Only on 422: one entry per offending field, with `loc`, `msg` and
            `type`.
          title: Details
        message:
          description: >-
            Human-readable explanation; safe to show to a developer, not meant
            for parsing.
          title: Message
          type: string
        trace_id:
          anyOf:
            - type: string
            - type: 'null'
          description: Quote this when contacting support; it identifies the exact request.
          title: Trace Id
      required:
        - code
        - message
      title: ErrorBody
      type: object
  securitySchemes:
    ApiKey:
      bearerFormat: ek_live_…
      description: >-
        An organization API key created in Settings → Team, sent as
        `Authorization: Bearer <key>`.
      scheme: bearer
      type: http

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.