curl --request POST \
--url https://api.apimart.ai/v1/videos/generations \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '{
"model": "viduq4-preview",
"prompt": "A girl turns back and smiles, her long hair blowing in the wind as the camera slowly moves closer",
"image_urls": ["https://example.com/first-frame.png"],
"duration": 5,
"resolution": "1080p"
}'
import requests
response = requests.post(
"https://api.apimart.ai/v1/videos/generations",
headers={"Authorization": "Bearer <token>"},
json={
"model": "viduq4-preview",
"prompt": "A girl turns back and smiles, her long hair blowing in the wind as the camera slowly moves closer",
"image_urls": ["https://example.com/first-frame.png"],
"duration": 5,
"resolution": "1080p"
}
)
response.raise_for_status()
print(response.json())
const response = await fetch("https://api.apimart.ai/v1/videos/generations", {
method: "POST",
headers: {
Authorization: "Bearer <token>",
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "viduq4-preview",
prompt: "A girl turns back and smiles, her long hair blowing in the wind as the camera slowly moves closer",
image_urls: ["https://example.com/first-frame.png"],
duration: 5,
resolution: "1080p"
})
});
if (!response.ok) throw new Error(await response.text());
console.log(await response.json());
{
"code": 200,
"data": [
{
"status": "submitted",
"task_id": "task_01K..."
}
]
}
Vidu Q4 Preview
Vidu Q4 Preview Video Generation
Generate videos from one first frame or up to 15 reference images and 3 reference audio clips. 3–16 seconds, up to 4K, with audio by default.
POST
/
v1
/
videos
/
generations
curl --request POST \
--url https://api.apimart.ai/v1/videos/generations \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '{
"model": "viduq4-preview",
"prompt": "A girl turns back and smiles, her long hair blowing in the wind as the camera slowly moves closer",
"image_urls": ["https://example.com/first-frame.png"],
"duration": 5,
"resolution": "1080p"
}'
import requests
response = requests.post(
"https://api.apimart.ai/v1/videos/generations",
headers={"Authorization": "Bearer <token>"},
json={
"model": "viduq4-preview",
"prompt": "A girl turns back and smiles, her long hair blowing in the wind as the camera slowly moves closer",
"image_urls": ["https://example.com/first-frame.png"],
"duration": 5,
"resolution": "1080p"
}
)
response.raise_for_status()
print(response.json())
const response = await fetch("https://api.apimart.ai/v1/videos/generations", {
method: "POST",
headers: {
Authorization: "Bearer <token>",
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "viduq4-preview",
prompt: "A girl turns back and smiles, her long hair blowing in the wind as the camera slowly moves closer",
image_urls: ["https://example.com/first-frame.png"],
duration: 5,
resolution: "1080p"
})
});
if (!response.ok) throw new Error(await response.text());
console.log(await response.json());
{
"code": 200,
"data": [
{
"status": "submitted",
"task_id": "task_01K..."
}
]
}
This model supports image-to-video and reference-to-video, but not text-only generation or first/last frames. After submission, read the task ID from
data[0].task_id and use Task Query to retrieve the status and result.Generation modes
The same model,viduq4-preview, automatically selects the mode based on images, roles, and reference audio. No additional mode parameter is needed.
| Input | Mode |
|---|---|
Only first_frame_image or one image with role: "first_frame" | Image-to-video |
| One image without a role and no reference audio | Image-to-video |
Images include a reference_image or reference role, with no explicit first frame | Reference-to-video |
| 2–15 images in total, with no explicit first frame | Reference-to-video |
| Reference audio with 1–15 images and no explicit first frame | Reference-to-video |
- Image-to-video: Exactly one first frame; prompt optional; reference audio is not accepted.
- Reference-to-video: 1–15 reference images, up to 3 reference audio clips, and a required prompt. With only one image and no reference audio, explicitly set
role: "reference_image"; otherwise, image-to-video is used. - An explicit first frame (
first_frame_imageorrole: "first_frame") cannot be combined with other images, reference image roles, or reference audio. Such combinations return HTTP 400.
curl --request POST \
--url https://api.apimart.ai/v1/videos/generations \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '{
"model": "viduq4-preview",
"prompt": "A girl turns back and smiles, her long hair blowing in the wind as the camera slowly moves closer",
"image_urls": ["https://example.com/first-frame.png"],
"duration": 5,
"resolution": "1080p"
}'
import requests
response = requests.post(
"https://api.apimart.ai/v1/videos/generations",
headers={"Authorization": "Bearer <token>"},
json={
"model": "viduq4-preview",
"prompt": "A girl turns back and smiles, her long hair blowing in the wind as the camera slowly moves closer",
"image_urls": ["https://example.com/first-frame.png"],
"duration": 5,
"resolution": "1080p"
}
)
response.raise_for_status()
print(response.json())
const response = await fetch("https://api.apimart.ai/v1/videos/generations", {
method: "POST",
headers: {
Authorization: "Bearer <token>",
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "viduq4-preview",
prompt: "A girl turns back and smiles, her long hair blowing in the wind as the camera slowly moves closer",
image_urls: ["https://example.com/first-frame.png"],
duration: 5,
resolution: "1080p"
})
});
if (!response.ok) throw new Error(await response.text());
console.log(await response.json());
{
"code": 200,
"data": [
{
"status": "submitted",
"task_id": "task_01K..."
}
]
}
Request headers
string
required
Bearer authentication in the format
Bearer <token>, where <token> is your APIMart API Key.Request parameters
string
required
Must be exactly
viduq4-preview, in lowercase.string
Video generation prompt, up to 20,000 characters.
- Image-to-video: Optional. If omitted, the model generates content based on the first frame.
- Reference-to-video: Required. If missing, returns HTTP 400.
string[]
Image array. Supports publicly accessible image URLs or Base64 Data URLs, such as
data:image/png;base64,....- Image-to-video: Exactly one image, used as the first frame.
- Reference-to-video: 1–15 images combined with
image_with_roles.
image_with_roles; counts are added together. Do not combine with first_frame_image or an explicit first_frame role. For a single image without a role, reference audio also determines whether reference-to-video is used.object[]
Array of images with roles. Use one element for image-to-video; for reference-to-video, 1–15 images combined with
May be combined with
image_urls.Show Show image fields
Show Show image fields
string
required
Public image URL or Base64 Data URL.
string
Image role, case-insensitive:
first_frame: First frame for image-to-video.reference_image: Reference image for reference-to-video;referenceis also accepted.- Omitted or empty: Without reference audio, the total count determines the mode: one image for image-to-video, two or more for reference-to-video. With reference audio, reference-to-video is used.
last_frame, synchronously return HTTP 400.image_urls to supply reference images, but first-frame roles cannot be mixed with reference assets.string
For image-to-video only. Provide a public URL or Base64 Data URL for the first frame.When using this field, do not provide other images or reference audio. For reference-to-video, use
image_urls or image_with_roles.string[]
Array of reference audio URLs, for reference-to-video only. At most 3 clips combined with
audio_url.MP3 required, 3–12 seconds per clip, up to 50MB each. Reference audio still requires at least one image and a prompt.Invalid audio format or duration causes the task to fail during execution with a full refund, rather than a synchronous HTTP 400 at submission.string
Single reference audio URL, with the same requirements as
audio_urls. At most 3 clips across both fields.string
default:"16:9"
For reference-to-video only. Supports
1:1, 9:16, 16:9, 3:4, and 4:3; defaults to 16:9.For image-to-video, the first frame determines the aspect ratio and this parameter is ignored.string
Compatibility alias for
aspect_ratio with the same allowed values. Use only one of these fields. Has no effect on image-to-video.integer
default:"5"
Video duration in seconds. Supports 3–16 seconds, not 1–2 seconds.
string
default:"720p"
Video resolution:
540p, 720p, 1080p, 2K, or 4K, case-insensitive.boolean
default:"true"
Whether to output video with dialogue and sound effects.
true: Video with an audio track (default).false: Silent video.
integer
Random seed. Omit or pass
0 for a random value.Asset requirements
- Image-to-video: Exactly one first-frame image is required; reference audio is not accepted.
- Reference-to-video: 1–15 reference images are required; up to 3 reference audio clips are optional.
- Supports PNG, JPEG, JPG, and WEBP, up to 50MB per image.
- With Base64, the entire request body must be under 20MB. Public URLs are recommended.
- Image URLs must be publicly accessible. Replace example URLs with actual accessible image URLs.
Both modes require images and do not support
last_frame_image. Parameter errors such as mixing first frames with reference assets or exceeding image/audio counts return HTTP 400 at submission, without creating a task or charging. Invalid reference audio format or duration causes failure during execution and a refund.Request examples
First frame only, without a prompt
{
"model": "viduq4-preview",
"image_urls": ["https://example.com/first-frame.png"]
}
First frame with an explicit role and 4K output
{
"model": "viduq4-preview",
"prompt": "The camera slowly moves closer as the person smiles naturally",
"image_with_roles": [
{
"url": "https://example.com/first-frame.png",
"role": "first_frame"
}
],
"duration": 8,
"resolution": "4K",
"audio": true
}
Silent video using the first-frame field
{
"model": "viduq4-preview",
"first_frame_image": "https://example.com/first-frame.png",
"duration": 5,
"resolution": "1080p",
"audio": false
}
Video from multiple images and reference audio
{
"model": "viduq4-preview",
"prompt": "The boy in image 1 speaks to the girl in image 2 using the content of the reference audio, in the cafe from image 3",
"image_urls": [
"https://example.com/boy.png",
"https://example.com/girl.png",
"https://example.com/cafe.png"
],
"audio_urls": ["https://example.com/line.mp3"],
"aspect_ratio": "16:9",
"duration": 8,
"resolution": "720p"
}
Reference-to-video with a single image
{
"model": "viduq4-preview",
"prompt": "The person in the reference image enters a cafe and waves to the staff",
"image_with_roles": [
{
"url": "https://example.com/person.png",
"role": "reference_image"
}
],
"aspect_ratio": "9:16",
"duration": 5,
"resolution": "1080p"
}
reference_image role. Replace all example image and audio URLs with actual accessible asset URLs.
Submission response
integer
Response status code;
200 indicates success.array
Query task results
Poll every 5–10 seconds and stop when the status iscompleted or failed. Use the unified query endpoint:
curl --request GET \
--url https://api.apimart.ai/v1/tasks/task_01K... \
--header 'Authorization: Bearer <token>'
{
"code": 200,
"data": {
"id": "task_01K...",
"status": "completed",
"progress": 100,
"result": {
"videos": [
{
"url": ["https://example.com/generated-video.mp4"]
}
]
}
}
}
| Status | Action |
|---|---|
pending | Queued; continue polling |
processing | Generating; continue polling |
completed | Success; read video links from the data.result.videos[0].url array |
failed | Failure; read the reason from data.error.message, stop polling, and receive a full refund |
status to determine completion, not fixed progress milestones.
Billing
Billed by video duration and resolution: cost = duration (seconds) × the per-second rate for the resolution. See Model Pricing for current prices. Image-to-video and reference-to-video cost the same, with or without audio. Reference images and audio incur no additional charge. Failed tasks are automatically fully refunded.Common parameter errors
The following synchronously return HTTP 400 without creating a task or charging:| Issue | Action |
|---|---|
| No images provided | Provide one first frame for image-to-video or 1–15 reference images for reference-to-video |
| Explicit first frame mixed with other images, reference roles, or reference audio | Keep only one first frame for image-to-video; remove explicit first-frame fields or roles for reference-to-video |
Unsupported role such as last_frame | Use first_frame, reference_image, reference, or leave empty |
| More than 15 reference images | Reduce the combined count of image_urls and image_with_roles to at most 15 |
| More than 3 reference audio clips | Limit audio_urls and audio_url to 3 clips combined |
Missing prompt in reference-to-video | Add a prompt of up to 20,000 characters |
Unsupported reference-to-video aspect ratio, such as 21:9 | Use 1:1, 9:16, 16:9, 3:4, or 4:3 |
last_frame_image provided | Remove this field; first/last frames are not supported |
duration below 3 or above 16 | Use an integer from 3 to 16 seconds |
Unsupported resolution such as 480p or 8K | Use 540p, 720p, 1080p, 2K, or 4K |
Other Vidu models
For text-to-video or first/last frames, use Vidu Q3 Pro / Turbo. This model already supports multiple reference images; Vidu Q3 Mix / Standard also offers reference-to-video. For 1–2 second clips, chooseviduq3-pro; this model requires at least 3 seconds.