AI video model
OmniHuman 1.5 AI Video Generator
Choose OmniHuman 1.5 when the project begins with one usable portrait image and one driving audio file, and the clip should follow those supplied media rather than a text-to-video workflow.
OmniHuman 1.5 task directions
These local references help turn the current OmniHuman 1.5 input and control contract into a concrete task direction. They are not presented as output from this model.
Every card identifies whether it is a verified model-family example or a visual reference. Reference cards are not model benchmarks or before-and-after claims.
- Input contract
- One image and one audio file are required
- Current controls
- Prompt length: Up to 1,000 characters · Duration: Follows the supplied audio (no duration selector) · Output resolution: 720P / 1080P · Framing: Framing follows the supplied media
- Selection boundary
- Do not use this workflow when you only have text or an image without audio. It does not make a promise about lip sync, identity preservation, or performance; compare a text- or image-led model when either required source file is unavailable.
Action beat
State one subject action and the moment the clip should begin from.
- Task direction
- State one subject action and the moment the clip should begin from.
Reference direction, not a model output
Local ImageStyle visual reference. It is not presented as a model-specific result.
Camera direction
Name the camera behavior and the visual change it should reveal over time.
- Task direction
- Name the camera behavior and the visual change it should reveal over time.
Reference direction, not a model output
Local ImageStyle visual reference. It is not presented as a model-specific result.
Ending state
Define the final visual state, without treating this reference as a result from the current model.
- Task direction
- Define the final visual state, without treating this reference as a result from the current model.
Reference direction, not a model output
Local ImageStyle visual reference. It is not presented as a model-specific result.
- Create
- AI video
- Input
- One image and one audio file are required
- Duration
- Follows the supplied audio (no duration selector)
Compare available settings with related models
Check whether the input, output, and framing controls fit the job before you open a workspace. This table does not treat the example media as a cross-model performance test.
OmniHuman 1.5
Current page
Portrait-led video that combines a supplied image with supplied audio.
- Input
- One image and one audio file are required
- Prompt length
- Up to 1,000 characters
- Duration
- Follows the supplied audio (no duration selector)
Seedance 2 Mini
Short multimodal clips with text, media references, audio generation, and web search.
- Input
- Text prompt or compatible reference media: image (up to 3), video (up to 3), audio (up to 3)
- Prompt length
- Up to 20,000 characters
- Duration
- 4–15 seconds
Seedance 2 Fast
The Fast-labelled multimodal workflow with a 15-second default clip.
- Input
- Text prompt or compatible reference media: image (up to 5), video (up to 3), audio (up to 3)
- Prompt length
- Up to 20,000 characters
- Duration
- 4–15 seconds
Seedance 2
Multimodal video work with explicit 480p through 4K output choices.
- Input
- Text prompt or compatible reference media: image (up to 5), video (up to 3), audio (up to 3)
- Prompt length
- Up to 20,000 characters
- Duration
- 4–15 seconds
What you can control with OmniHuman 1.5
Build from portrait and audio together
OmniHuman 1.5 currently requires one portrait image and one audio file. The supplied audio determines duration, supplied media determines framing, and the workspace exposes 720p, 1080p, and fast mode.
Upload the required media
One image and one audio file are required. Only upload media you are allowed to use for the current project.
Confirm the clip settings
Current clip settings include Follows the supplied audio (no duration selector), 720P / 1080P, and Framing follows the supplied media.
Current OmniHuman 1.5 settings
These settings come from the current ImageStyle model configuration and help you decide whether the workflow fits the job.
| Setting | Current availability |
|---|---|
| Input | One image and one audio file are required |
| Prompt length | Up to 1,000 characters |
| Duration | Follows the supplied audio (no duration selector) |
| Output resolution | 720P / 1080P |
| Framing | Framing follows the supplied media |
| Additional controls | Current core settings |
AI video model
When to choose OmniHuman 1.5
Portrait-led video that combines a supplied image with supplied audio.
Current workflow fit
Choose OmniHuman 1.5 when the project begins with one usable portrait image and one driving audio file, and the clip should follow those supplied media rather than a text-to-video workflow.
When to compare another model
Do not use this workflow when you only have text or an image without audio. It does not make a promise about lip sync, identity preservation, or performance; compare a text- or image-led model when either required source file is unavailable.
How to create a video with OmniHuman 1.5
Prepare the portrait and audio as the two source-of-truth inputs, then use the optional prompt only to clarify the requested treatment around those supplied media.
- 1
Verify the two required sources
Confirm that the portrait subject and driving audio are authorized for the project and belong together before setting the output option.
- 2
Check image and audio alignment
Keep the uploaded subject aligned with the language, timing, or performance in the audio. Add a written direction only when it helps clarify the result.
- 3
Confirm duration and output controls
Confirm the current Follows the supplied audio (no duration selector), 720P / 1080P, and Framing follows the supplied media in the workspace before generation.
- 4
Generate and review continuity
Before public use, inspect subject motion, camera changes, visible text, sound, and the final frame.
Checks before you create and publish
- Treat the workspace as current
- Model availability and settings can change. Use the options shown in the current workspace before you submit.
- Use approved source media
- When you upload images, video, or audio, confirm that you are allowed to use and publish it for the current project.
- Review the generated video
- Before publishing, manually check motion continuity, text, sound, people, product details, and marks.
OmniHuman 1.5 FAQ
- Can I use OmniHuman 1.5 with text only or an image only?
- No. The current workflow requires one portrait image and one audio file. Use a text- or image-led video model when either source is unavailable.
- Can I start OmniHuman 1.5 from reference media?
- One image and one audio file are required. Check the current media type and quantity limits before generation, and upload only approved source material.
- How long can a OmniHuman 1.5 video be?
- The current video-duration setup is Follows the supplied audio (no duration selector).
- What output controls does OmniHuman 1.5 currently offer?
- Current output options include 720P / 1080P and Framing follows the supplied media. Other available controls include current core settings.
- Can I publish a generated result without review?
- No. Before public use, manually review motion continuity, text, sound, product details, and any other element that must be accurate.
Updated:
Capability source: Current ImageStyle video-model configuration.