Skip to main content
Google’s Gemini text-to-speech model. Converts text to natural-sounding speech using Gemini’s advanced voice synthesis.

Request

Required: text
Optional: voice_name
Reference inputs take asset ids from POST /v1/uploads or an earlier output’s asset_id. Prices and allowed values come from the live catalog; GET /v1/models/gemini-3.1-flash-tts-preview is always current.