Skip to main content
POST
Use this endpoint to start a Grok video job from text, a source image, or reference inputs. It returns a request_id immediately, so treat it as the first step in an async workflow. Always send model: grok-imagine-video-1.5 explicitly. model and prompt are required for the 1.5 request shapes documented here.

Choose an input mode

For a data URI, use data:image/png;base64,<BASE64_IMAGE_DATA>. Each reference_images item uses the same url field. A reference audio item uses a preset voice_id, such as eve. In the prompt, tags such as <IMAGE_0> and <AUDIO_0> can identify the corresponding reference.

Start with a small request

  • Use model: grok-imagine-video-1.5
  • For a first request, keep duration at 1 and resolution at 480p
  • Keep prompt explicit so the scene and motion are clear
  • Add one input mode at a time before combining reference images and voices

Duration and resolution

Generated videos include an audio track. Set duration, resolution, and aspect_ratio explicitly when your application depends on a particular output shape.

Task flow

1

Create the job

Send the model, prompt, and any inputs, then save the returned request_id.
2

Poll for completion

Call Get xAI video results until the top-level status becomes done, failed, or expired. An initial response can contain only the echoed request_id while task metadata is prepared.
3

Persist the output

Copy the final video.url into your own storage if you need it after the temporary delivery window.

Authorizations

Authorization
string
header
required

Use your CometAPI API key as the bearer value.

Body

application/json
model
enum<string>
required

The Grok video model ID. Send this field explicitly so the request uses Grok Imagine Video 1.5.

Available options:
grok-imagine-video-1.5
Example:

"grok-imagine-video-1.5"

prompt
string
required

A description of the scene, motion, and audio. Reference-to-video prompts can identify inputs with tags such as <IMAGE_0> and <AUDIO_0>.

Minimum string length: 1
Example:

"A paper boat glides across a quiet pond at sunrise."

duration
integer
default:8

The output duration in seconds. Use an integer from 1 through 15.

Required range: 1 <= x <= 15
aspect_ratio
enum<string>

The output aspect ratio. Text-to-video uses 16:9 when omitted. Image-to-video uses the source image's aspect ratio when omitted.

Available options:
1:1,
16:9,
9:16,
4:3,
3:4,
3:2,
2:3
Example:

"16:9"

resolution
enum<string>
default:480p

The output resolution. Reference-to-video requests support 480p and 720p; text-to-video and image-to-video also support 1080p.

Available options:
480p,
720p,
1080p
image
object

The source image for image-to-video. Do not combine this field with reference_images or reference_audios.

reference_images
object[]

Visual references for reference-to-video. Each item accepts a public image URL or image data URI. Do not combine this field with image.

Required array length: 1 - 7 elements
reference_audios
object[]

Preset voice references for reference-to-video. Do not combine this field with image.

Required array length: 1 - 3 elements

Response

200 - application/json

Request accepted.

request_id
string
required

The deferred request ID used to poll GET /grok/v1/videos/{request_id}.

Last modified on August 21, 2026