AI APIs for Developers
Filter by use case
AI Model APIs
106 modelsAlibaba
Wan 3.0 Reference To Video
Feed Wan 3.0 one or more reference images and generate video that keeps faces, characters, and objects consistent shot after shot — ideal for branded content, avatars, and multi-scene storytelling via the ModelsLab API.
Alibaba
Wan 3.0 Image To Video
Bring any photo to life with Wan 3.0's image-to-video model. Upload a source image, describe the motion, and get a smooth, physically consistent AI video — built for product, portrait, and creative animation via the ModelsLab API.
Alibaba
Wan 3.0 Text To Video
Turn any text prompt into a cinematic AI video with Wan 3.0 — Alibaba's next-gen video model. Sharper motion, longer clips, and stronger prompt adherence than Wan 2.5, available now via the ModelsLab API.
ModelsLab
MiniMax H3 Start/ End Frame
MiniMax H3 is an open-weight multimodal AI video generation model capable of creating up to 15-second videos with native stereo audio from text, images, video, and audio inputs. It supports text-to-video, image-to-video, video editing, and video-to-video
ModelsLab
Minimax H3 Reference to Video
MiniMax H3 is an open-weight multimodal AI video generation model capable of creating up to 15-second videos with native stereo audio from text, images, video, and audio inputs. It supports text-to-video, image-to-video, video editing, and video-to-video
ltx
LTX 2.5 Pro Text To Video
LTX 2.5 Pro turns text prompts into cinematic video with native synchronized audio — dialogue, ambience, and sound effects generated alongside the frames. Output at 720p or 1080p, 25 or 50 fps, up to 10 seconds, from a single ModelsLab API call.
ltx
LTX 2.5 Pro Image To Video
LTX 2.5 Pro Image to Video animates a single still into a cinematic clip, using a text prompt to direct motion, camera movement, and atmosphere. Native synchronized audio, 720p or 1080p output, 25 or 50 fps, and 6 to 10 second durations — all from one end
ModelsLab
Minimax H3 Text to Video
MiniMax H3 is an open-weight multimodal AI video generation model capable of creating up to 15-second videos with native stereo audio from text, images, video, and audio inputs. It supports text-to-video, image-to-video, video editing, and video-to-video
Bytedance
Seedance 2.5 Text to Video
Seedance 2.5 writes a full scene from one text prompt — up to 30 seconds of continuous video with synced dialogue, ambience and score, no stitching required. Prompt in 11 languages, pick any aspect ratio from 21:9 to 9:16, and render at 480p or 720p
Bytedance
Seedance 2.5 Image To Video
Seedance 2.5 animates a single still into a video up to 30 seconds long, with native sound and dialogue in 11 languages. Set a first frame, or pin both first and last frames to control exactly where the shot starts and ends. Outputs 480p or 720p.
Bytedance
Seedance 2.5 Multimodal Reference to Video
Seedance 2.5 turns up to 50 multimodal references 30 images, 10 video clips and 10 audio tracks into one coherent video up to 30 seconds long, with native audio in 11 languages. Lock characters, products, motion and sound in a single call at 480p or 720p.
Black Forest Labs
Flux 3 Video To Video
FLUX 3 in video-continuation mode. Upload an MP4 and FLUX 3 carries the shot on from its final frames for another 5-20 seconds, audio included.
Black Forest Labs
Flux 3 Image To Video
FLUX 3 in image-to-video mode. Drop in 1-10 images as keyframes, set the timing, and get a 5-20s HD clip with synchronized audio.
Black Forest Labs
Flux 3 Text To Video
Black Forest Labs' FLUX 3 in text-to-video mode. Type a prompt, get a 5-20s HD or FHD clip with synchronized audio built in
Minmax
MiniMax H3 Start/ End Frame To Video
Give H3 a start frame and an end frame — it generates the transition in native 2K with sound. Deterministic in and out points for edit-ready clips.
Minmax
MiniMax H3 Image To Video
Feed one image and a prompt, get a native 2K clip with sound. Preserves subject identity, lighting, and on-image text across motion.















