Skip to main content
Google’s high-speed multimodal video model for native-audio generation, first/last-frame interpolation, subject references, uploaded-video editing, and conversational refinement.

Text → Video

Generate a native-audio video from a text prompt. Select it with "mode": "text_to_video", or leave mode out and the inputs decide. Required: prompt
Optional: aspect_ratio
Resolution tiers: 360p, 720p, 1080p, 4k (send as tier).

Frames → Video

Animate a first frame, optionally interpolating to a last frame. Select it with "mode": "image_to_video", or leave mode out and the inputs decide. Required: prompt, first_frame
Optional: last_frame, aspect_ratio
Resolution tiers: 360p, 720p, 1080p, 4k (send as tier).

References → Video

Use up to 10 images as subject and visual references. Select it with "mode": "reference_to_video", or leave mode out and the inputs decide. Required: prompt, reference_images
Optional: aspect_ratio
Resolution tiers: 360p, 720p, 1080p, 4k (send as tier).

Edit Video

Apply a prompt-driven edit to an uploaded 3–10 second video. Select it with "mode": "video_edit", or leave mode out and the inputs decide. Required: prompt, source_video
Optional: none
Resolution tiers: 360p, 720p, 1080p, 4k (send as tier).
Reference inputs take asset ids from POST /v1/uploads or an earlier output’s asset_id. Prices and allowed values come from the live catalog; GET /v1/models/gemini-omni-1.1-flash is always current.