AI video model

OmniHuman 1.5 AI Video Generator

Choose OmniHuman 1.5 when the project begins with one usable portrait image and one driving audio file, and the clip should follow those supplied media rather than a text-to-video workflow.

OmniHuman 1.5 task directions

These local references help turn the current OmniHuman 1.5 input and control contract into a concrete task direction. They are not presented as output from this model.

Every card identifies whether it is a verified model-family example or a visual reference. Reference cards are not model benchmarks or before-and-after claims.

Input contract
One image and one audio file are required
Current controls
Prompt length: Up to 1,000 characters · Duration: Follows the supplied audio (no duration selector) · Output resolution: 720P / 1080P · Framing: Framing follows the supplied media
Selection boundary
Do not use this workflow when you only have text or an image without audio. It does not make a promise about lip sync, identity preservation, or performance; compare a text- or image-led model when either required source file is unavailable.

Action beat

State one subject action and the moment the clip should begin from.

Task direction
State one subject action and the moment the clip should begin from.

Reference direction, not a model output

Local ImageStyle visual reference. It is not presented as a model-specific result.

Create with OmniHuman 1.5

Camera direction

Name the camera behavior and the visual change it should reveal over time.

Task direction
Name the camera behavior and the visual change it should reveal over time.

Reference direction, not a model output

Local ImageStyle visual reference. It is not presented as a model-specific result.

Create with OmniHuman 1.5

Ending state

Define the final visual state, without treating this reference as a result from the current model.

Task direction
Define the final visual state, without treating this reference as a result from the current model.

Reference direction, not a model output

Local ImageStyle visual reference. It is not presented as a model-specific result.

Create with OmniHuman 1.5
Create
AI video
Input
One image and one audio file are required
Duration
Follows the supplied audio (no duration selector)

Compare available settings with related models

Check whether the input, output, and framing controls fit the job before you open a workspace. This table does not treat the example media as a cross-model performance test.

What you can control with OmniHuman 1.5

Build from portrait and audio together

OmniHuman 1.5 currently requires one portrait image and one audio file. The supplied audio determines duration, supplied media determines framing, and the workspace exposes 720p, 1080p, and fast mode.

Upload the required media

One image and one audio file are required. Only upload media you are allowed to use for the current project.

Confirm the clip settings

Current clip settings include Follows the supplied audio (no duration selector), 720P / 1080P, and Framing follows the supplied media.

Current OmniHuman 1.5 settings

These settings come from the current ImageStyle model configuration and help you decide whether the workflow fits the job.

SettingCurrent availability
InputOne image and one audio file are required
Prompt lengthUp to 1,000 characters
DurationFollows the supplied audio (no duration selector)
Output resolution720P / 1080P
FramingFraming follows the supplied media
Additional controlsCurrent core settings

AI video model

When to choose OmniHuman 1.5

Portrait-led video that combines a supplied image with supplied audio.

Current workflow fit

Choose OmniHuman 1.5 when the project begins with one usable portrait image and one driving audio file, and the clip should follow those supplied media rather than a text-to-video workflow.

When to compare another model

Do not use this workflow when you only have text or an image without audio. It does not make a promise about lip sync, identity preservation, or performance; compare a text- or image-led model when either required source file is unavailable.

How to create a video with OmniHuman 1.5

Prepare the portrait and audio as the two source-of-truth inputs, then use the optional prompt only to clarify the requested treatment around those supplied media.

  1. 1

    Verify the two required sources

    Confirm that the portrait subject and driving audio are authorized for the project and belong together before setting the output option.

  2. 2

    Check image and audio alignment

    Keep the uploaded subject aligned with the language, timing, or performance in the audio. Add a written direction only when it helps clarify the result.

  3. 3

    Confirm duration and output controls

    Confirm the current Follows the supplied audio (no duration selector), 720P / 1080P, and Framing follows the supplied media in the workspace before generation.

  4. 4

    Generate and review continuity

    Before public use, inspect subject motion, camera changes, visible text, sound, and the final frame.

Checks before you create and publish

Treat the workspace as current
Model availability and settings can change. Use the options shown in the current workspace before you submit.
Use approved source media
When you upload images, video, or audio, confirm that you are allowed to use and publish it for the current project.
Review the generated video
Before publishing, manually check motion continuity, text, sound, people, product details, and marks.

OmniHuman 1.5 FAQ

Can I use OmniHuman 1.5 with text only or an image only?
No. The current workflow requires one portrait image and one audio file. Use a text- or image-led video model when either source is unavailable.
Can I start OmniHuman 1.5 from reference media?
One image and one audio file are required. Check the current media type and quantity limits before generation, and upload only approved source material.
How long can a OmniHuman 1.5 video be?
The current video-duration setup is Follows the supplied audio (no duration selector).
What output controls does OmniHuman 1.5 currently offer?
Current output options include 720P / 1080P and Framing follows the supplied media. Other available controls include current core settings.
Can I publish a generated result without review?
No. Before public use, manually review motion continuity, text, sound, product details, and any other element that must be accurate.

Updated:

Capability source: Current ImageStyle video-model configuration.

Create with OmniHuman 1.5