AI video model
Grok Imagine Video AI Video Generator
Choose Grok Imagine Video when you need a Grok AI video generator that starts from one image or a written direction and the 6 to 30 second duration range fits the intended narrative beat.
Grok Imagine Video task directions
These local references help turn the current Grok Imagine Video input and control contract into a concrete task direction. They are not presented as output from this model.
Every card identifies whether it is a verified model-family example or a visual reference. Reference cards are not model benchmarks or before-and-after claims.
- Input contract
- Text prompt or compatible reference media: image (up to 1)
- Current controls
- Prompt length: Up to 5,000 characters · Duration: 6–30 seconds · Output resolution: 480P / 720P · Framing: Text-to-video: 5 aspect-ratio choices; image-to-video: follows the supplied image
- Selection boundary
- Choose Grok Imagine Video 1.5 Preview when the task needs up to three image references, a 1 to 15 second range, or an adaptive frame option. In this model’s image-to-video mode, framing follows the supplied image rather than a manual selector.
Action beat
State one subject action and the moment the clip should begin from.
- Task direction
- State one subject action and the moment the clip should begin from.
Reference direction, not a model output
Local ImageStyle visual reference. It is not presented as a model-specific result.
Camera direction
Name the camera behavior and the visual change it should reveal over time.
- Task direction
- Name the camera behavior and the visual change it should reveal over time.
Reference direction, not a model output
Local ImageStyle visual reference. It is not presented as a model-specific result.
Ending state
Define the final visual state, without treating this reference as a result from the current model.
- Task direction
- Define the final visual state, without treating this reference as a result from the current model.
Reference direction, not a model output
Local ImageStyle visual reference. It is not presented as a model-specific result.
- Create
- AI video
- Input
- Text prompt or compatible reference media: image (up to 1)
- Duration
- 6–30 seconds
Compare available settings with related models
Check whether the input, output, and framing controls fit the job before you open a workspace. This table does not treat the example media as a cross-model performance test.
Grok Imagine Video
Current page
Text or single-image video with 6 to 30 seconds and image-led framing.
- Input
- Text prompt or compatible reference media: image (up to 1)
- Prompt length
- Up to 5,000 characters
- Duration
- 6–30 seconds
Grok Imagine Video 1.5 Preview
Preview text or multi-image video workflow with 1 to 15 seconds and adaptive framing.
- Input
- Text prompt or compatible reference media: image (up to 3)
- Prompt length
- Up to 4,096 characters
- Duration
- 1–15 seconds
Seedance 2 Mini
Short multimodal clips with text, media references, audio generation, and web search.
- Input
- Text prompt or compatible reference media: image (up to 3), video (up to 3), audio (up to 3)
- Prompt length
- Up to 20,000 characters
- Duration
- 4–15 seconds
Seedance 2 Fast
The Fast-labelled multimodal workflow with a 15-second default clip.
- Input
- Text prompt or compatible reference media: image (up to 5), video (up to 3), audio (up to 3)
- Prompt length
- Up to 20,000 characters
- Duration
- 4–15 seconds
What you can control with Grok Imagine Video
Start a longer single-image or text direction
Grok Imagine Video currently supports text or one image, 480p or 720p, and 6 to 30 second clips. An image-led clip uses the source image for framing.
Add compatible media when it helps
Text prompt or compatible reference media: image (up to 1). Only upload media you are allowed to use for the current project.
Confirm the clip settings
Current clip settings include 6–30 seconds, 480P / 720P, and Text-to-video: 5 aspect-ratio choices; image-to-video: follows the supplied image.
Current Grok Imagine Video settings
These settings come from the current ImageStyle model configuration and help you decide whether the workflow fits the job.
| Setting | Current availability |
|---|---|
| Input | Text prompt or compatible reference media: image (up to 1) |
| Prompt length | Up to 5,000 characters |
| Duration | 6–30 seconds |
| Output resolution | 480P / 720P |
| Framing | Text-to-video: 5 aspect-ratio choices; image-to-video: follows the supplied image |
| Additional controls | Current core settings |
AI video model
When to choose Grok Imagine Video
Text or single-image video with 6 to 30 seconds and image-led framing.
Current workflow fit
Choose Grok Imagine Video when you need a Grok AI video generator that starts from one image or a written direction and the 6 to 30 second duration range fits the intended narrative beat.
When to compare another model
Choose Grok Imagine Video 1.5 Preview when the task needs up to three image references, a 1 to 15 second range, or an adaptive frame option. In this model’s image-to-video mode, framing follows the supplied image rather than a manual selector.
How to create a video with Grok Imagine Video
For image-led work, choose one source image whose existing frame already fits the delivery; for text-led work, write the scene and motion around the selected duration.
- 1
Choose the duration before the direction
Select the narrative beat within 6 to 30 seconds, then write one main action or choose the single source image that establishes the scene.
- 2
Add reference media when needed
Add approved compatible media only when an existing frame, motion, or sound should guide the result.
- 3
Confirm duration and output controls
Confirm the current 6–30 seconds, 480P / 720P, and Text-to-video: 5 aspect-ratio choices; image-to-video: follows the supplied image in the workspace before generation.
- 4
Generate and review continuity
Before public use, inspect subject motion, camera changes, visible text, sound, and the final frame.
Checks before you create and publish
- Treat the workspace as current
- Model availability and settings can change. Use the options shown in the current workspace before you submit.
- Use approved source media
- When you upload images, video, or audio, confirm that you are allowed to use and publish it for the current project.
- Review the generated video
- Before publishing, manually check motion continuity, text, sound, people, product details, and marks.
Grok Imagine Video FAQ
- What are Grok Imagine video generation capabilities in ImageStyle?
- The current xAI Grok Imagine video generation workflow accepts text or one source image, supports 6 to 30 second clips, and uses the source image to determine framing for image-led generation.
- Can Grok Imagine Video start from an image without a prompt?
- Yes. The current single-image workflow permits an image-only start. The supplied image determines framing in that mode.
- Can I start Grok Imagine Video from reference media?
- Text prompt or compatible reference media: image (up to 1). Check the current media type and quantity limits before generation, and upload only approved source material.
- How long can a Grok Imagine Video video be?
- The current video-duration setup is 6–30 seconds.
- What output controls does Grok Imagine Video currently offer?
- Current output options include 480P / 720P and Text-to-video: 5 aspect-ratio choices; image-to-video: follows the supplied image. Other available controls include current core settings.
- Can I publish a generated result without review?
- No. Before public use, manually review motion continuity, text, sound, product details, and any other element that must be accurate.
Updated:
Capability source: Current ImageStyle video-model configuration.