Model directory
Choose an AI image or video model from the job you need to finish.
Each model page explains the current input contract, available output controls, and practical workflow inside ImageStyle. Open a model when you know the result you need, then verify the settings before generating.
AI image models
Compare text-led and reference-guided image workflows, framing choices, and output tiers.
Nano Banana 2
Reference-guided images with manual 1K, 2K, or 4K selection and broad framing.
View model detailsNano Banana 2 Lite
Reference-guided image directions with broad framing and provider-selected output.
View model detailsNano Banana Pro
Image work with up to eight approved references and manual 1K, 2K, or 4K selection.
View model detailsNano Banana
Text-only image direction with flexible framing and provider-selected output.
View model detailsSeedream 5 Pro
Focused text- or reference-guided image work with 1K or 2K output choices.
View model detailsGPT Image 2
Detailed briefs and references with a 1K, 2K, or 4K framing trade-off.
View model detailsGrok Imagine
A text or single-reference image direction with Standard or High mode selection.
View model detailsGrok Imagine Image 2.0
Text-to-image with a 5,000-character prompt limit and five documented frames.
View model detailsZ Image
Concise text-to-image work with five fixed frames and provider-selected output.
View model detailsWan 2.7 Image
Reference-guided images with 1K or 2K output and eight fixed frames.
View model detailsWan 2.7 Image Pro
Text-led image output with 1K, 2K, or 4K; 4K is unavailable with references.
View model detailsQwen Image 3.0
Text-to-image and reference-guided image editing with 1K or 2K output and eight fixed frames.
View model detailsQwen Image 3.0 Pro
Higher-fidelity text-to-image and reference-guided image editing with 1K or 2K output.
View model detailsAI video models
Compare text, image, audio, and motion-guided video workflows, duration, framing, and output options.
Seedance 2 Mini
Short multimodal clips with text, media references, audio generation, and web search.
View model detailsSeedance 2 Fast
The Fast-labelled multimodal workflow with a 15-second default clip.
View model detailsSeedance 2
Multimodal video work with explicit 480p through 4K output choices.
View model detailsSeedance 2.5
Longer multimodal video with 480p or 720p output, ordered frames, and media references.
View model detailsHappyHorse 1.1
Text or one-start-image clips with nine current text-to-video frames.
View model detailsHappyHorse
A concise text or one-first-frame workflow with five standard frames.
View model detailsKling 2.6
Prompt-led 5 or 10 second clips with fixed 720p and audio generation.
View model detailsKling V3 Turbo
Prompt-led text or image video with 720p or 1080p and 3 to 15 seconds.
View model detailsKling 3.0 Video
Prompt-led short sequences with one image, multi-shot, audio, and quality tiers.
View model detailsKling O3
Single-shot text or ordered-frame video with 720p to 4K output and optional native audio.
View model detailsWan 2.6 Video
Prompt-guided transformations of supplied source videos with 5 or 10 second output, 720p or 1080p, and multi-shot control.
View model detailsKling 2.6 Motion Control
Reference-image and driving-video motion control with 720p or 1080p, and orientation-dependent source duration.
View model detailsKling 3.0 Motion Control
Reference-image and driving-video motion control with 720p or 1080p, and orientation-dependent source duration.
View model detailsWan 2.7 Video
Prompt-led video direction with frames, motion clip, driving audio, and negative prompt.
View model detailsGrok Imagine Video
Text or single-image video with 6 to 30 seconds and image-led framing.
View model detailsGrok Imagine Video 1.5 Preview
Preview text or multi-image video workflow with 1 to 15 seconds and adaptive framing.
View model detailsGemini Omni Video
Long prompt-led video briefs with optional images and fixed 4, 6, 8, or 10 seconds.
View model detailsMiniMax H3
Text, start/end-frame, or media-reference video with 768P or 2K output and 4 to 15 second clips.
View model detailsVeo 3.1 Quality
The Quality-labelled text or first-and-last-frame workflow for 4, 6, or 8 seconds.
View model detailsVeo 3.1 Fast
The Fast-labelled text or first-and-last-frame workflow for 4, 6, or 8 seconds.
View model detailsVeo 3.1 Lite
The Lite-labelled text or first-and-last-frame workflow for 4, 6, or 8 seconds.
View model detailsOmniHuman 1.5
Portrait-led video that combines a supplied image with supplied audio.
View model detailsModel availability and settings can change. The generation workspace shows the current controls before you submit.
Capability source: Current ImageStyle model configuration.