# Ekly API > Generate images, video, music and speech from your own code, with the same models and credits as the Ekly app. - [Introduction](https://docs.ekly.ai/introduction.md): What the Ekly API does, what it does not do yet, and who it is for. - [Quickstart](https://docs.ekly.ai/quickstart.md): Create a key, pick a model, generate an image and download it. - [Authentication](https://docs.ekly.ai/authentication.md): Organization API keys, where they come from, and what they can do. - [Generations](https://docs.ekly.ai/generations.md): The lifecycle of a generation, how to poll, and how idempotency keys protect you from double charges. - [Uploads](https://docs.ekly.ai/uploads.md): Register a reference image, video or audio file and get the asset id a generation can use. - [Credits and limits](https://docs.ekly.ai/credits-and-limits.md): How the API spends credits, how to see a price first, and what a per-member cap means for a key. - [Errors](https://docs.ekly.ai/errors.md): Every error has the same shape and a small, stable set of codes. - [Rate limits](https://docs.ekly.ai/rate-limits.md): What the limits are, how to read the headers, and how to behave on a 429. - [Versioning](https://docs.ekly.ai/versioning.md): What can change under /v1 and what cannot. - [Changelog](https://docs.ekly.ai/changelog.md): What changed in the Ekly API, newest first. - [Images](https://docs.ekly.ai/guides/images.md): Which image model to reach for, and the fields that matter. - [Video](https://docs.ekly.ai/guides/video.md): Text-to-video, image-to-video, reference video and extensions. - [Speech and music](https://docs.ekly.ai/guides/audio.md): One endpoint, two kinds of model. - [Custom Avatar Generation](https://docs.ekly.ai/models/avatar-custom-flux.md): Generate a custom avatar face image using FLUX Schnell. - [FLUX 2 Pro](https://docs.ekly.ai/models/flux-2-pro-fal.md): FLUX 2 Pro — next-generation image model with improved quality, prompt adherence, and multi-reference editing. - [FLUX Kontext Pro](https://docs.ekly.ai/models/flux-kontext-pro-fal.md): Image editing with character consistency. - [FLUX Pro Ultra](https://docs.ekly.ai/models/flux-2-flex-fal.md): High-resolution image generation with raw photographic output. - [GPT Image 1.5](https://docs.ekly.ai/models/gpt-image-1-5.md): OpenAI GPT Image 1.5 via the direct OpenAI Images API. - [GPT Image 2](https://docs.ekly.ai/models/gpt-image-2.md): OpenAI GPT Image 2 via the direct OpenAI Images API. - [GPT Image 2.5 Flare](https://docs.ekly.ai/models/gpt-image-2-5-flare.md): Fast everyday image generation and reference editing. - [GPT Image 2.5 Sunburst](https://docs.ekly.ai/models/gpt-image-2-5-sunburst.md): Detailed image generation and precise reference-image editing. - [Grok Imagine](https://docs.ekly.ai/models/grok-imagine-image-fal.md): xAI's image generation model powered by the Aurora engine. - [Grok Imagine Image 2.0](https://docs.ekly.ai/models/grok-imagine-image-2-0.md): Create polished images or edit up to three references with Grok Imagine Image 2.0. - [Ideogram Character](https://docs.ekly.ai/models/ideogram-character-fal.md): Ideogram character consistency model — generate images with consistent character identity from reference images. - [Ideogram V3](https://docs.ekly.ai/models/ideogram-v3-fal.md): Advanced image generation with best-in-class text rendering in images. - [Nano Banana](https://docs.ekly.ai/models/gemini-2-5-flash-image.md): Latest Gemini model with streaming image generation. - [Nano Banana Flash](https://docs.ekly.ai/models/gemini-3-1-flash-image.md): Gemini 3.1 Flash with Pro-level image quality at Flash-tier speed and pricing. - [Nano Banana Pro](https://docs.ekly.ai/models/gemini-3-pro-image.md): Advanced Gemini 3 Pro model with 4K image generation, real-world grounding, and thinking process. - [Qwen Image Edit Sassy](https://docs.ekly.ai/models/qwen-image-edit-sassy.md): Change one thing in a picture you already have. - [Recraft V3](https://docs.ekly.ai/models/recraft-v3-fal.md): State-of-the-art image generation with excellent text rendering and multiple artistic styles. - [Recraft V4](https://docs.ekly.ai/models/recraft-v4-fal.md): Professional-grade image generation with lifelike realism. - [Seedream 5.0 Lite](https://docs.ekly.ai/models/seedream-5-0-lite-byteplus.md): BytePlus Seedream 5.0 Lite image generation with text-to-image, reference image editing, and sequential image generation. - [Seedream 5.0 Pro](https://docs.ekly.ai/models/seedream-5-0-pro-byteplus.md): BytePlus Seedream 5.0 Pro high-precision image generation with text-to-image and multi-reference editing (up to 10 images). - [Z-Image Sassy](https://docs.ekly.ai/models/z-image-sassy.md): Quick text-to-image on the Custom lane. - [Creatify Aurora](https://docs.ekly.ai/models/creatify-aurora-fal.md): Studio-quality talking or singing avatar video from a face photo and audio. - [FLUX.3 Video](https://docs.ekly.ai/models/flux-3-fal.md): Black Forest Labs FLUX.3 video generation with text, image, first/last-frame, exact keyframe, and video-extension modes with native audio. - [Gemini Omni 1.1 Flash](https://docs.ekly.ai/models/gemini-omni-1-1-flash.md): Google's high-speed multimodal video model for native-audio generation, first/last-frame interpolation, subject references, uploaded-video editing, and conversational refinement. - [Gemini Omni Flash](https://docs.ekly.ai/models/gemini-omni-flash-preview.md): Google's conversational video model (preview): text/image-to-video with native audio, plus natural-language refinement of a previous generation. - [Gemini Veo 3.1](https://docs.ekly.ai/models/veo-3-1-generate-001.md): Google's Veo 3.1 video generation model with support for up to 3 reference images. - [Gemini Veo 3.1 Fast](https://docs.ekly.ai/models/veo-3-1-fast-generate-001.md): Google's Veo 3.1 Fast video generation model with support for start/end frames and reference images. - [Gemini Veo 3.1 Lite](https://docs.ekly.ai/models/veo-3-1-lite-generate-001.md): Lightweight Veo 3.1 — fastest and cheapest with frames support (reference images not supported on Lite). - [Genjutsu Motion Transfer](https://docs.ekly.ai/models/genjutsu-motion-transfer.md): Keep the motion and timing of a clip you already have, and re-perform it with the characters, products, or clothes in your reference images. - [Genjutsu Object Swap](https://docs.ekly.ai/models/genjutsu-object-swap.md): Swap an object in a clip you already have for the one in your reference images — a product, a prop, a garment — while the rest of the shot stays as filmed. - [Grok Imagine Video](https://docs.ekly.ai/models/grok-imagine-video-fal.md): xAI's video generation with native audio, dialogue, and sound effects in a single pass. - [Grok Imagine Video 1.5](https://docs.ekly.ai/models/grok-imagine-video-1-5-fal.md): Generate videos with native synchronized audio from a prompt, an opening image, or up to seven reference images. - [Grok Imagine Video Edit](https://docs.ekly.ai/models/grok-imagine-video-edit-fal.md): xAI's video-to-video editor. - [Kling AI Video Pro](https://docs.ekly.ai/models/kling-v1-6-fal.md): Kling's advanced video generation model with multiple modes. - [Kling O3 Omni](https://docs.ekly.ai/models/kling-o3-fal.md): Kling O3 Omni — native-audio video with multi-shot, element references, reference-to-video, restyle-from-video and prompt-based editing. - [Kling V3](https://docs.ekly.ai/models/kling-v3-fal.md): Kling V3 — text & image to video with first/last frame, multi-shot and element references. - [Kling V3 Motion Control](https://docs.ekly.ai/models/kling-v2-6-motion-control-fal.md): Transfer real human motion, gestures, and expressions from a reference video to a character image using Kling V3. - [LTX Video 2.3](https://docs.ekly.ai/models/ltx-2-3-fal.md): Fast, high-quality video generation up to 4K resolution. - [LTX Video 2.3 Audio-to-Video](https://docs.ekly.ai/models/ltx-2-3-audio-fal.md): Generate video driven by an audio clip (2-20s) with LTX-2.3 — lip-sync and motion follow the sound. - [LTX Video 2.3 Extend](https://docs.ekly.ai/models/ltx-2-3-extend-fal.md): Extend an existing clip with newly generated, continuous footage and native audio using LTX-2.3. - [LTX Video 2.3 Fast](https://docs.ekly.ai/models/ltx-2-3-fast-fal.md): Speed-optimized LTX-2.3 (~30x faster) for quick text-to-video and image-to-video with start/end frame control and native audio. - [Luma Ray 2](https://docs.ekly.ai/models/luma-ray-2-fal.md): Luma's Ray 2 — cinematic video generation with natural motion and physics understanding. - [MiniMax H3](https://docs.ekly.ai/models/minimax-h3-fal.md): MiniMax H3 video generation with text, first/last-frame animation, and multimodal reference guidance at 480p, 768p, 2K, or 4K. - [MiniMax H3 Max](https://docs.ekly.ai/models/minimax-h3-max-fal.md): fal's post-trained MiniMax H3 variant for stronger prompt adherence, polished aesthetics, and text, first/last-frame, or multimodal reference video generation. - [MiniMax Hailuo 2.3](https://docs.ekly.ai/models/minimax-hailuo-2-3-fal.md): MiniMax's latest Hailuo 2.3 model with text-to-video and image-to-video. - [P-Video 2 Pro](https://docs.ekly.ai/models/p-video-2-pro.md): Pruna's highest-quality video model. - [PixVerse V5](https://docs.ekly.ai/models/pixverse-v5-fal.md): Versatile image-to-video with artistic styles (anime, 3D, comic, cyberpunk). - [Seedance 1.5 Pro](https://docs.ekly.ai/models/seedance-1-5-pro-fal.md): ByteDance's latest video generation with native audio, 4-12 second output, up to 1080p. - [Seedance 2.0](https://docs.ekly.ai/models/seedance-2-0.md): ByteDance Seedance 2.0 through fal Enterprise with native audio, first/last-frame control, and multimodal image, video, and audio references. - [Seedance 2.0 Fast](https://docs.ekly.ai/models/seedance-2-0-fast.md): Fast Seedance 2.0 through fal Enterprise with native audio, first/last-frame control, and multimodal references. - [Seedance 2.0 Mini](https://docs.ekly.ai/models/seedance-2-0-mini.md): Fastest, most affordable Seedance 2.0 tier through fal with native audio, first/last-frame control, and multimodal references. - [Seedance 2.5](https://docs.ekly.ai/models/seedance-2-5-fal.md): ByteDance Seedance 2.5 generates coherent continuous videos or multi-shot sequences with native audio, strong prompt and layout adherence, first/last-frame control, and up to 50 multimodal references. - [Sora 2 (OpenAI)](https://docs.ekly.ai/models/sora-2-fal.md): OpenAI's Sora 2 video generation. - [Sync Lipsync 3](https://docs.ekly.ai/models/sync-lipsync-v3-fal.md): Re-sync the lips in a video you already have to a new audio track. - [Topaz Precision](https://docs.ekly.ai/models/topaz-precision-fal.md): Sharpen and upscale real footage without inventing anything. - [Topaz Starlight](https://docs.ekly.ai/models/topaz-starlight-fal.md): Rebuild detail that isn't in your clip anymore. - [Topaz Starlight Fast](https://docs.ekly.ai/models/topaz-starlight-fast-fal.md): Topaz's fastest generative upscaler at half the price of Starlight. - [VEED Fabric](https://docs.ekly.ai/models/veed-fabric-fal.md): Turn a face photo and an audio clip into a lip-synced talking video. - [Vidu Q3](https://docs.ekly.ai/models/vidu-q3-fal.md): Vidu Q3 — affordable, fast video generation with text-to-video and image-to-video. - [Wan 2.2 InfiniteTalk](https://docs.ekly.ai/models/wan-2-2-infinitetalk-fal.md): Generates realistic talking avatar videos from a face image and audio. - [Wan 2.6](https://docs.ekly.ai/models/wan-2-6-fal.md): Alibaba's Wan 2.6 — high-quality text-to-video and image-to-video at excellent value. - [Wan 3.0](https://docs.ekly.ai/models/wan-3-0-fal.md): Alibaba Wan 3.0 generates coherent videos from text, first and optional last frames, or multimodal image, video, and audio references with native audio. - [Wan 3.0 Prime](https://docs.ekly.ai/models/wan-3-0-prime-fal.md): Alibaba Wan 3.0 Prime generates production-focused videos from text, first and optional last frames, or multimodal image, video, and audio references with native audio. - [ElevenLabs Multilingual](https://docs.ekly.ai/models/eleven-multilingual-v2.md): High-quality multilingual TTS supporting 29 languages with natural intonation and emotion. - [ElevenLabs Music v1](https://docs.ekly.ai/models/music-v1.md): Studio-grade text-to-music from ElevenLabs. - [ElevenLabs Music v2](https://docs.ekly.ai/models/music-v2.md): ElevenLabs Music v2 text-to-music with stronger prompt adherence and richer vocals. - [ElevenLabs Turbo](https://docs.ekly.ai/models/eleven-turbo-v2-5.md): Low-latency TTS model optimized for real-time applications with high quality output. - [ElevenLabs v3](https://docs.ekly.ai/models/eleven-v3.md): Most expressive TTS with 70+ languages, dramatic delivery, and support for audio tags (e.g. - [ElevenLabs v4](https://docs.ekly.ai/models/eleven-v4.md): ElevenLabs' most expressive speech model: 90+ languages, stronger voice likeness, and natural-language audio tags (e.g. - [Gemini TTS](https://docs.ekly.ai/models/gemini-3-1-flash-tts-preview.md): Google's Gemini text-to-speech model. - [Lyria 2 (High Quality)](https://docs.ekly.ai/models/lyria-002.md): Google's production Lyria 2 model via Vertex AI. - [Lyria 3](https://docs.ekly.ai/models/lyria-3-clip-preview.md): Google's Lyria 3 model for 30-second high-fidelity music clips. - [Lyria 3 Pro](https://docs.ekly.ai/models/lyria-3-pro-preview.md): Google's Lyria 3 Pro model for full-length compositions up to 2 minutes. - [Lyria RealTime](https://docs.ekly.ai/models/models-lyria-realtime-exp.md): Google's real-time AI music generation model. - [Qwen 3 Text-to-Speech](https://docs.ekly.ai/models/fal-qwen3-tts-1-7b.md): High-quality multilingual TTS from Alibaba's Qwen 3 model via fal.ai. - [Qwen 3 Voice Design](https://docs.ekly.ai/models/fal-qwen3-voice-design-1-7b.md): Design a custom voice style via text prompt — describe the tone, accent, and emotion you want. - [Sarvam Bulbul v2](https://docs.ekly.ai/models/sarvam-bulbul-v2.md): Indic-native TTS (v2) with 8 prebuilt voices across 11 Indian languages. - [Sarvam Bulbul v3](https://docs.ekly.ai/models/sarvam-bulbul-v3.md): Indic-native TTS with 8 prebuilt voices across 11 Indian languages. - [MCP server](https://docs.ekly.ai/mcp.md): Use the same generation tools from Claude, ChatGPT, Cursor and other agents without writing code. - [Generate speech or music](https://docs.ekly.ai/api-reference/v1/generate-speech-or-music.md): Submits a speech or music generation; the model decides which. Returns immediately with `status: pending`; credits are reserved now and settled when the job finishes. Poll `GET /v1/generations/{id}` until `completed` or `failed`. - [Get credit balance](https://docs.ekly.ai/api-reference/v1/get-credit-balance.md): Credits available to this key right now. Generations charge the organization's balance as the key's creator, so a per-member cap on that person applies here too. - [List generations](https://docs.ekly.ai/api-reference/v1/list-generations.md): Generations created by this key's creator in this organization, newest first. Use `next_cursor` as `after` to page; cursors stay valid as new rows arrive. - [Get a generation](https://docs.ekly.ai/api-reference/v1/get-a-generation.md): The current state of one generation. Poll every 2–5 seconds with backoff; images usually finish in under a minute, video in a few minutes. Output URLs appear when `status` is `completed`. - [Generate images](https://docs.ekly.ai/api-reference/v1/generate-images.md): Submits an image generation. Reference inputs such as `reference_images` take asset ids from `POST /v1/uploads` or an earlier output. Returns immediately with `status: pending`; credits are reserved now and settled when the job finishes. Poll `GET /v1/generations/{id}` until `completed` or `failed`. - [List models](https://docs.ekly.ai/api-reference/v1/list-models.md): Every generation model available to your organization, with the parameters each accepts and its credit price. Call this first; the `id` is what you pass as `model` when generating. - [Get a model](https://docs.ekly.ai/api-reference/v1/get-a-model.md): Parameters, modes, tiers, defaults, input rules and pricing for one model, exactly as the catalog declares them. Use it to build a valid request before you estimate or generate. - [Estimate cost](https://docs.ekly.ai/api-reference/v1/estimate-cost.md): The credits a generation would charge for these settings, after your organization's discounts, and whether your balance and plan allow it. An estimate reserves nothing; the actual charge happens when you generate. - [Create an upload](https://docs.ekly.ai/api-reference/v1/create-an-upload.md): Registers a reference file and returns a signed URL for its bytes. PUT the file to `upload_url` with exactly the `required_headers`, then pass `asset_id` in a generation request (for example as `reference_images` or `first_frame`). The upload URL is valid for ten minutes; the asset appears in the wo… - [Generate a video](https://docs.ekly.ai/api-reference/v1/generate-a-video.md): Submits a video generation from a prompt, a first/last frame, reference media, an earlier generation (`source_generation_id`, to extend it) or a character image plus motion video (motion-control models). `GET /v1/models/{id}` lists the inputs each mode needs; reference inputs take asset ids. Returns… ## OpenAPI Specs - [openapi.v1](/openapi.v1.json) This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.