Skip to main content
Models that take references (image-to-video, first and last frames, motion control, lip sync, voice reference) take asset ids. The upload flow gives you one for a file on your side, and every generation output already has one, so chaining generations needs no upload at all.
1

Register the file

The response has asset_id (what you will pass to a generation), upload_url (where to PUT the bytes) and required_headers (send them exactly). The size you declare is binding; a PUT with a different number of bytes is rejected. For video and audio, send duration_seconds too.
2

PUT the file

Send the required_headers from the previous response exactly as given; the declared size is part of the signature, so a file of a different size is refused. Upload URLs expire after ten minutes. Register the file again if you miss it.
3

Use the asset id in a generation

GET /v1/models/{id} names the reference inputs each model and mode accept (first_frame, reference_images, motion_video, …) and which are required.

Accepted files

Images (PNG, JPEG, WebP), video (MP4, MOV, WebM) and audio (MP3, WAV, M4A). Per-model limits on count, duration and frame rate come from GET /v1/models/{id}; a request that breaks them is rejected before any credit is reserved. An asset id must belong to your organization and be visible to the key’s creator, the same rule the app’s Assets library applies. Uploaded files appear in that library like any other upload.