Skip to main content
fal’s post-trained MiniMax H3 variant for stronger prompt adherence, polished aesthetics, and text, first/last-frame, or multimodal reference video generation.

Text → Video

Generate a video from a prompt. Select it with "mode": "text_to_video", or leave mode out and the inputs decide. Required: prompt, duration_seconds
Optional: aspect_ratio
Resolution tiers: 480P, 768P (send as tier).

Image → Video

Animate a first frame, with an optional last frame. Output framing follows the first image. Select it with "mode": "image_to_video", or leave mode out and the inputs decide. Required: prompt, first_frame, duration_seconds
Optional: last_frame
Resolution tiers: 480P, 768P (send as tier).

Reference → Video

Attach at least one reference image or video to guide subjects, style, motion, and sound. Audio needs a visual reference; up to 12 files are accepted. Select it with "mode": "reference_to_video", or leave mode out and the inputs decide. Required: prompt, duration_seconds
At least one of: reference_images, reference_videos
Optional: reference_audios, aspect_ratio
Resolution tiers: 480P, 768P (send as tier).
Reference inputs take asset ids from POST /v1/uploads or an earlier output’s asset_id. Prices and allowed values come from the live catalog; GET /v1/models/minimax-h3-max-fal is always current.