Vidu Video Generation API
This guide explains how to call Vidu video models through this service. Use an API token issued by this service, and replace BASE_URL and API_TOKEN in the examples with your actual values.
Advanced parameters are passed through to the upstream service via metadata. Use PascalCase field names, such as Audio, AudioType, CallbackUrl, and LogoAdd.
Endpoint Overview
The OpenAI Video-compatible endpoints are recommended:
| Method | Path | Purpose | Response style |
|---|---|---|---|
POST | /v1/videos | Create a video generation task | OpenAI Video object |
GET | /v1/videos/{task_id} | Query task status and result | OpenAI Video object |
GET | /v1/videos/{task_id}/content | Download the generated video | Video binary |
Legacy endpoints are also supported:
| Method | Path | Purpose |
|---|---|---|
POST | /v1/video/generations | Create a video generation task |
GET | /v1/video/generations/{task_id} | Query task status and result |
POST | /v1/videos/generations | Create a video generation task |
Capability Routing
This service automatically selects the upstream Action based on the input:
| Request shape | Submit Action | Query Action | Capability |
|---|---|---|---|
No image / images / input_reference | SubmitTextToVideoViduJob | DescribeTextToVideoViduJob | Text-to-video |
Includes image or 1 to 2 images | SubmitImageToVideoViduJob | DescribeImageToVideoViduJob | Image-to-video; 2 images are treated as start and end frames |
Includes input_reference, 3 or more images, or metadata.Action=referenceGenerate / metadata.action=referenceGenerate | SubmitReferenceToVideoViduJob | DescribeReferenceToVideoViduJob | Reference-to-video |
metadata.Action / metadata.action is only used by this service for routing and is not forwarded upstream.
Supported Models
| API model name | Upstream Model | Text-to-video | Image-to-video | Reference-to-video | Notes |
|---|---|---|---|---|---|
viduq3-pro | viduq3-pro | Supported | Supported | Not recommended | Recommended model for text-to-video and image-to-video |
viduq3-turbo | viduq3-turbo | Supported | Supported | Not recommended | Faster generation than viduq3-pro |
viduq3-pro-fast | viduq3-pro-fast | Not recommended | Requires upstream support | Not recommended | Recognized by this service; requires actual enablement on the upstream account |
viduq2-pro | viduq2-pro | Not supported | Supported | Submitted as viduq2 | Image-to-video model; this service downgrades it to viduq2 for reference-to-video |
viduq2-turbo | viduq2-turbo | Not supported | Supported | Submitted as viduq2 | Image-to-video model |
viduq2-pro-fast | viduq2-pro-fast | Not supported | Requires upstream support | Submitted as viduq2 | Recognized by this service; requires actual enablement on the upstream account |
viduq2 | viduq2 | Supported | Not supported | Supported | Text-to-video and reference-to-video model |
For production calls, prefer the model names that are explicitly marked as supported above. Use fast suffix models only when the upstream account has actually enabled them.
Authentication
All requests use a Bearer token:
Authorization: Bearer sk-...
Content-Type: application/jsonTop-level Request Fields
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | Yes | None | API model name from the table above |
prompt | string | Required for text-to-video and reference-to-video; optional for image-to-video | None | Video prompt. Upstream limit: no more than 2000 characters |
image | string | Available for image-to-video | None | Shorthand for a single input image; when provided, the request is handled as image-to-video |
images | string[] | Available for image-to-video or reference-to-video | None | 1 image creates first-frame image-to-video, 2 images create start/end-frame image-to-video, and 3 to 7 images create reference-to-video |
input_reference | string | No | None | Triggers reference-to-video; this service uses it as reference image input |
duration | integer/string | No | 5 | Video duration in seconds. See the model matrix below for available ranges |
metadata | object | No | {} | Extension parameters using PascalCase field names |
Images can be public URLs. Vidu image-to-video and reference-to-video Actions only accept URL strings. If Base64 is provided, this service first converts it to a temporary URL when COS is configured; otherwise the upstream service may return an invalid image URL error.
Model Parameter Matrix
Duration
| Capability | Model | Available duration range |
|---|---|---|
| Text-to-video | viduq3-pro, viduq3-turbo | Integers from 1 to 16, default 5 |
| Text-to-video | viduq2 | Integers from 1 to 10, default 5 |
| First-frame image-to-video | viduq3-pro, viduq3-turbo | Integers from 1 to 16, default 5 |
| First-frame image-to-video | viduq2-pro, viduq2-turbo | Integers from 1 to 10, default 5 |
| Start/end-frame image-to-video | viduq3-pro, viduq3-turbo | Integers from 1 to 16, default 5 |
| Start/end-frame image-to-video | viduq2-pro, viduq2-turbo | Integers from 1 to 8, default 5 |
| Reference-to-video | viduq2 | Integers from 1 to 10, default 5 |
Resolution and Aspect Ratio
| Parameter | Available values | Service behavior |
|---|---|---|
AspectRatio | 16:9, 9:16, 4:3, 3:4, 1:1; reference subject calls commonly use 16:9, 9:16, 1:1 | Passed through via metadata.AspectRatio |
Resolution | 540p, 720p, 1080p, usually defaulting to 720p | This service currently filters metadata.Resolution / metadata.resolution and does not forward it |
MovementAmplitude | auto, small, medium, large | This service currently filters this field; it also does not take effect for q2/q3 series models |
Style | general, anime | Does not take effect for q2/q3 series models |
metadata Parameters
Do not pass Model, Prompt, Images, or Duration in metadata. These fields are generated from top-level model, prompt, image / images, and duration; duplicating them can cause mismatched request content and billing model.
| Field | Type | Applicable capability | Description |
|---|---|---|---|
AspectRatio | string | Text-to-video, reference-to-video | Output aspect ratio, commonly 16:9, 9:16, 4:3, 3:4, 1:1 |
Bgm | boolean | Text-to-video, image-to-video, reference-to-video | Whether to add system preset background music. It does not take effect for Q3 series models; for Q2 series models, it does not take effect at 9 or 10 seconds |
Audio | boolean | Text-to-video, image-to-video, reference-to-video | Whether to use direct audio/video output. Text-to-video supports it only on Q3 series models; reference-to-video supports it only for subject calls |
AudioType | string | Direct audio/video output | Can be passed when Audio=true; common values are all, speech_only, sound_effect_only |
VoiceId | string | Image-to-video or reference subject | Specifies the voice. If empty, the system recommends one. Voice cloning is not currently supported |
IsRec | boolean | Image-to-video | Whether to use the system-recommended prompt. When enabled, the upstream service does not use the top-level prompt, and additional tokens are consumed |
Subjects | array<object> | Reference-to-video | Subject reference information. See "Reference Subjects" |
Videos | string[] | Reference-to-video | Video reference, supported only by viduq2-pro; passed through via metadata.Videos |
MetaData | string | All capabilities | Metadata identifier as a JSON string. If empty, Vidu default metadata is used |
CallbackUrl | string | All capabilities | Upstream callback URL. The callback does not replace task querying through this service |
Payload | string | All capabilities | Pass-through parameter, up to 1048576 characters |
OffPeak | boolean | All capabilities | Off-peak mode. When enabled, it consumes fewer tokens, but the task may complete within 48 hours; unfinished tasks are canceled and refunded as tokens |
LogoAdd | integer | All capabilities | Watermark switch: 1 adds a watermark, 0 does not add one, and other values are treated as 1 |
LogoParam | object | All capabilities | Custom watermark. See "Watermark Parameters" for fields |
Vidu uses LogoAdd / LogoParam to control explicit watermarks. It does not use Watermark, WmPosition, or WmUrl.
Request Examples
Text-to-video
curl -X POST "${BASE_URL}/v1/videos" \
-H "Authorization: Bearer ${API_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"model": "viduq3-turbo",
"prompt": "Santa Claus and a bear hug beside a lake, snow falls slowly, slight camera push-in, warm cinematic lighting",
"duration": 5,
"metadata": {
"AspectRatio": "16:9",
"Audio": true
}
}'First-frame Image-to-video
curl -X POST "${BASE_URL}/v1/videos" \
-H "Authorization: Bearer ${API_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"model": "viduq3-pro",
"prompt": "The person naturally raises their head and looks at the camera, background lights flow subtly, keep the face clear",
"image": "https://cdn.example.com/first-frame.png",
"duration": 5,
"metadata": {
"Audio": true,
"AudioType": "all"
}
}'Start/end-frame Image-to-video
When 2 images are provided, the first image is used as the start frame and the second as the end frame. The resolutions of the two images should be close; the recommended start-frame resolution / end-frame resolution ratio is between 0.8 and 1.25.
{
"model": "viduq3-turbo",
"prompt": "Transition naturally from the start frame to the end frame with continuous motion and smooth camera movement",
"images": [
"https://cdn.example.com/start.png",
"https://cdn.example.com/end.png"
],
"duration": 5,
"metadata": {
"Audio": true
}
}Reference-to-video
Passing 3 to 7 images automatically uses SubmitReferenceToVideoViduJob. You can also use metadata.Action=referenceGenerate to force reference-to-video.
curl -X POST "${BASE_URL}/v1/videos" \
-H "Authorization: Bearer ${API_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"model": "viduq2",
"prompt": "Preserve the appearance of the person in the reference images while they walk naturally on a futuristic city street",
"images": [
"https://cdn.example.com/character-front.png",
"https://cdn.example.com/character-side.png",
"https://cdn.example.com/scene.png"
],
"duration": 5,
"metadata": {
"AspectRatio": "16:9"
}
}'Reference Subjects
Subjects is used for subject calls. It supports 1 to 7 subjects, with 1 to 7 total subject images; each subject can have at most 3 images. You can reference subjects in the prompt with @subject_id.
{
"model": "viduq2",
"prompt": "@subject_1 and @subject_2 talk at a street-side cafe, with a voiceover saying the weather is nice today",
"duration": 5,
"metadata": {
"Action": "referenceGenerate",
"Subjects": [
{
"Id": "subject_1",
"Name": "subject_1",
"Images": [
"https://cdn.example.com/person-a-front.png",
"https://cdn.example.com/person-a-side.png"
],
"VoiceId": "male-qn-qingse"
},
{
"Id": "subject_2",
"Name": "subject_2",
"Images": [
"https://cdn.example.com/person-b-front.png"
]
}
],
"Audio": true,
"AudioType": "speech_only"
}
}Reference subject fields:
| Field | Type | Required | Description |
|---|---|---|---|
Id | string | Yes | Subject ID, referenced in the prompt as @Id |
Images | string[] | Yes | Subject image URLs; each subject can have up to 3 images |
Name | string | No | Subject name, usually the same as Id |
Videos | string[] | No | Subject video URLs; supported only by viduq2-pro |
VoiceId | string | No | Subject voice ID |
Video Reference
Video reference is passed through via metadata.Videos. This capability is supported only by viduq2-pro, with up to one 8-second video or two 5-second videos. Supported formats are mp4, avi, and mov; size must not exceed 100 MB.
{
"model": "viduq2",
"prompt": "Use the character's motion rhythm from the reference video to generate a subject-consistent video",
"duration": 5,
"metadata": {
"Action": "referenceGenerate",
"Videos": [
"https://cdn.example.com/reference-motion.mp4"
]
}
}Callback, Off-peak Mode, Watermark, and Business Payload
{
"model": "viduq3-pro",
"prompt": "A product commercial shot with the subject rotating slowly against soft background lighting",
"image": "https://cdn.example.com/product.png",
"duration": 5,
"metadata": {
"Payload": "client-order-20260909-0001",
"CallbackUrl": "https://example.com/callbacks/video",
"OffPeak": true,
"LogoAdd": 0
}
}The upstream callback does not replace task querying through this service. Clients should still poll status using the task_id returned by this service.
Input Asset Constraints
| Asset | Constraints |
|---|---|
| Image-to-video image | Supports URL or Base64; formats png, jpeg, jpg, webp; size no more than 50 MB; avoid aspect ratios over 1:4 or 4:1 |
| Start/end-frame images | Pass 2 images: the first is the start frame and the second is the end frame. The two image resolutions must be close, with start-frame resolution / end-frame resolution between 0.8 and 1.25 |
| Reference images | images supports 1 to 7 images; formats png, jpeg, jpg, webp; pixels at least 128x128; size no more than 50 MB |
| Reference subject images | Subjects supports 1 to 7 subjects, with 1 to 7 total subject images; each subject can have at most 3 images |
| Reference videos | Supports one 8-second video or two 5-second videos; formats mp4, avi, mov; pixels at least 128x128; size no more than 100 MB |
Watermark Parameters
{
"metadata": {
"LogoAdd": 1,
"LogoParam": {
"LogoUrl": "https://cdn.example.com/logo.png",
"LogoRect": {
"X": -222,
"Y": -54,
"Width": 202,
"Height": 34
}
}
}
}| Field | Type | Description |
|---|---|---|
LogoUrl | string | Watermark image URL |
LogoImage | string | Watermark image Base64; if both LogoImage and LogoUrl are provided, LogoUrl takes precedence |
LogoRect.X | integer | Watermark rectangle X coordinate; positive values move from left to right, negative values from right to left |
LogoRect.Y | integer | Watermark rectangle Y coordinate; positive values move from top to bottom, negative values from bottom to top |
LogoRect.Width | integer | Watermark rectangle width in px |
LogoRect.Height | integer | Watermark rectangle height in px |
Submission Response
After POST /v1/videos succeeds, the service returns a public task ID. The task runs asynchronously and must be queried later.
{
"id": "task_7a31e02c4b",
"task_id": "task_7a31e02c4b",
"object": "video",
"model": "viduq3-pro",
"status": "queued",
"progress": 0,
"created_at": 1788912000
}The upstream JobId in the original submission response is mapped by this service to the upstream ID of the internal task. Clients only need to save the task_id returned by this service.
Query a Task
curl "${BASE_URL}/v1/videos/task_7a31e02c4b" \
-H "Authorization: Bearer ${API_TOKEN}"Example response when completed:
{
"id": "task_7a31e02c4b",
"object": "video",
"model": "viduq3-pro",
"status": "completed",
"progress": 100,
"created_at": 1788912000,
"completed_at": 1788912060,
"metadata": {
"url": "https://example.com/generated-video.mp4"
}
}Upstream query statuses are mapped to this service's task statuses:
Upstream Status | Meaning | Service status |
|---|---|---|
WAIT | Waiting | queued / submitted |
RUN | Running | in_progress |
DONE | Task succeeded | completed |
FAIL | Task failed | failed |
The upstream ResultVideoUrl is valid for 24 hours. If COS persistence is configured, this service copies the upstream temporary URL to its own COS URL. If COS is not configured or persistence fails, the upstream temporary URL is returned.
Polling Example
TASK_ID="task_7a31e02c4b"
while true; do
BODY=$(curl -s "${BASE_URL}/v1/videos/${TASK_ID}" \
-H "Authorization: Bearer ${API_TOKEN}")
echo "${BODY}"
STATUS=$(printf '%s' "${BODY}" | jq -r '.status')
if [ "${STATUS}" = "completed" ] || [ "${STATUS}" = "failed" ]; then
break
fi
sleep 5
donePoll every 3 to 5 seconds. Do not create duplicate video tasks for the same business request.
Download the Video
curl -L "${BASE_URL}/v1/videos/task_7a31e02c4b/content" \
-H "Authorization: Bearer ${API_TOKEN}" \
-o output.mp4If metadata.url in the query response or result_url in the legacy response is a temporary URL, it may expire. For long-term retention, download the video and copy it to your own object storage.
Common Errors
| Symptom | Possible cause | Suggested handling |
|---|---|---|
unsupported model | The model name is not in this service's model list, or the administrator has not enabled the model | Check the model spelling and backend model configuration |
| Text-to-video request rejected | An image-to-video-only model was used, such as viduq2-pro or viduq2-turbo | Switch to viduq2, viduq3-pro, or viduq3-turbo |
| Image-to-video request rejected | A text-to-video-only model was used, or the image URL / Base64 is invalid | Switch to an image-to-video model; check image format, size, aspect ratio, and accessibility |
| Base64 image failed | Vidu Actions require URLs, COS conversion is not configured in this service, or the Base64 cannot be parsed | Use a public HTTPS image directly, or ask an administrator to configure COS |
| Reference-to-video failed | A non-viduq2 model was used, image/subject/video count is invalid, or asset format is invalid | Prefer viduq2; reduce input count or replace assets according to the asset constraints |
Resolution does not take effect | This service currently filters metadata.Resolution / metadata.resolution | Use the default resolution for now; adapter changes are required to support it |
Watermark does not take effect | Vidu uses LogoAdd / LogoParam | Use metadata.LogoAdd=0 or pass LogoParam |
| Audio not generated | The model or capability does not support Audio, or AudioType / VoiceId does not match | Prefer Q3 series for text-to-video; enable audio for reference-to-video only in subject calls |
| Query never completes | The upstream task is still in WAIT / RUN, or OffPeak placed it in an off-peak queue | Keep polling; off-peak mode may take up to 48 hours |
| Download failed | Upstream URL expired or was blocked by proxy security policy | Download promptly; ask an administrator to check COS persistence and the video proxy if needed |