Skip to content

Vidu Video Generation API ​

This guide explains how to call Vidu video models through this service. Use an API token issued by this service, and replace BASE_URL and API_TOKEN in the examples with your actual values.

Advanced parameters are passed through to the upstream service via metadata. Use PascalCase field names, such as Audio, AudioType, CallbackUrl, and LogoAdd.

Endpoint Overview ​

The OpenAI Video-compatible endpoints are recommended:

MethodPathPurposeResponse style
POST/v1/videosCreate a video generation taskOpenAI Video object
GET/v1/videos/{task_id}Query task status and resultOpenAI Video object
GET/v1/videos/{task_id}/contentDownload the generated videoVideo binary

Legacy endpoints are also supported:

MethodPathPurpose
POST/v1/video/generationsCreate a video generation task
GET/v1/video/generations/{task_id}Query task status and result
POST/v1/videos/generationsCreate a video generation task

Capability Routing ​

This service automatically selects the upstream Action based on the input:

Request shapeSubmit ActionQuery ActionCapability
No image / images / input_referenceSubmitTextToVideoViduJobDescribeTextToVideoViduJobText-to-video
Includes image or 1 to 2 imagesSubmitImageToVideoViduJobDescribeImageToVideoViduJobImage-to-video; 2 images are treated as start and end frames
Includes input_reference, 3 or more images, or metadata.Action=referenceGenerate / metadata.action=referenceGenerateSubmitReferenceToVideoViduJobDescribeReferenceToVideoViduJobReference-to-video

metadata.Action / metadata.action is only used by this service for routing and is not forwarded upstream.

Supported Models ​

API model nameUpstream ModelText-to-videoImage-to-videoReference-to-videoNotes
viduq3-providuq3-proSupportedSupportedNot recommendedRecommended model for text-to-video and image-to-video
viduq3-turboviduq3-turboSupportedSupportedNot recommendedFaster generation than viduq3-pro
viduq3-pro-fastviduq3-pro-fastNot recommendedRequires upstream supportNot recommendedRecognized by this service; requires actual enablement on the upstream account
viduq2-providuq2-proNot supportedSupportedSubmitted as viduq2Image-to-video model; this service downgrades it to viduq2 for reference-to-video
viduq2-turboviduq2-turboNot supportedSupportedSubmitted as viduq2Image-to-video model
viduq2-pro-fastviduq2-pro-fastNot supportedRequires upstream supportSubmitted as viduq2Recognized by this service; requires actual enablement on the upstream account
viduq2viduq2SupportedNot supportedSupportedText-to-video and reference-to-video model

For production calls, prefer the model names that are explicitly marked as supported above. Use fast suffix models only when the upstream account has actually enabled them.

Authentication ​

All requests use a Bearer token:

http
Authorization: Bearer sk-...
Content-Type: application/json

Top-level Request Fields ​

FieldTypeRequiredDefaultDescription
modelstringYesNoneAPI model name from the table above
promptstringRequired for text-to-video and reference-to-video; optional for image-to-videoNoneVideo prompt. Upstream limit: no more than 2000 characters
imagestringAvailable for image-to-videoNoneShorthand for a single input image; when provided, the request is handled as image-to-video
imagesstring[]Available for image-to-video or reference-to-videoNone1 image creates first-frame image-to-video, 2 images create start/end-frame image-to-video, and 3 to 7 images create reference-to-video
input_referencestringNoNoneTriggers reference-to-video; this service uses it as reference image input
durationinteger/stringNo5Video duration in seconds. See the model matrix below for available ranges
metadataobjectNo{}Extension parameters using PascalCase field names

Images can be public URLs. Vidu image-to-video and reference-to-video Actions only accept URL strings. If Base64 is provided, this service first converts it to a temporary URL when COS is configured; otherwise the upstream service may return an invalid image URL error.

Model Parameter Matrix ​

Duration ​

CapabilityModelAvailable duration range
Text-to-videoviduq3-pro, viduq3-turboIntegers from 1 to 16, default 5
Text-to-videoviduq2Integers from 1 to 10, default 5
First-frame image-to-videoviduq3-pro, viduq3-turboIntegers from 1 to 16, default 5
First-frame image-to-videoviduq2-pro, viduq2-turboIntegers from 1 to 10, default 5
Start/end-frame image-to-videoviduq3-pro, viduq3-turboIntegers from 1 to 16, default 5
Start/end-frame image-to-videoviduq2-pro, viduq2-turboIntegers from 1 to 8, default 5
Reference-to-videoviduq2Integers from 1 to 10, default 5

Resolution and Aspect Ratio ​

ParameterAvailable valuesService behavior
AspectRatio16:9, 9:16, 4:3, 3:4, 1:1; reference subject calls commonly use 16:9, 9:16, 1:1Passed through via metadata.AspectRatio
Resolution540p, 720p, 1080p, usually defaulting to 720pThis service currently filters metadata.Resolution / metadata.resolution and does not forward it
MovementAmplitudeauto, small, medium, largeThis service currently filters this field; it also does not take effect for q2/q3 series models
Stylegeneral, animeDoes not take effect for q2/q3 series models

metadata Parameters ​

Do not pass Model, Prompt, Images, or Duration in metadata. These fields are generated from top-level model, prompt, image / images, and duration; duplicating them can cause mismatched request content and billing model.

FieldTypeApplicable capabilityDescription
AspectRatiostringText-to-video, reference-to-videoOutput aspect ratio, commonly 16:9, 9:16, 4:3, 3:4, 1:1
BgmbooleanText-to-video, image-to-video, reference-to-videoWhether to add system preset background music. It does not take effect for Q3 series models; for Q2 series models, it does not take effect at 9 or 10 seconds
AudiobooleanText-to-video, image-to-video, reference-to-videoWhether to use direct audio/video output. Text-to-video supports it only on Q3 series models; reference-to-video supports it only for subject calls
AudioTypestringDirect audio/video outputCan be passed when Audio=true; common values are all, speech_only, sound_effect_only
VoiceIdstringImage-to-video or reference subjectSpecifies the voice. If empty, the system recommends one. Voice cloning is not currently supported
IsRecbooleanImage-to-videoWhether to use the system-recommended prompt. When enabled, the upstream service does not use the top-level prompt, and additional tokens are consumed
Subjectsarray<object>Reference-to-videoSubject reference information. See "Reference Subjects"
Videosstring[]Reference-to-videoVideo reference, supported only by viduq2-pro; passed through via metadata.Videos
MetaDatastringAll capabilitiesMetadata identifier as a JSON string. If empty, Vidu default metadata is used
CallbackUrlstringAll capabilitiesUpstream callback URL. The callback does not replace task querying through this service
PayloadstringAll capabilitiesPass-through parameter, up to 1048576 characters
OffPeakbooleanAll capabilitiesOff-peak mode. When enabled, it consumes fewer tokens, but the task may complete within 48 hours; unfinished tasks are canceled and refunded as tokens
LogoAddintegerAll capabilitiesWatermark switch: 1 adds a watermark, 0 does not add one, and other values are treated as 1
LogoParamobjectAll capabilitiesCustom watermark. See "Watermark Parameters" for fields

Vidu uses LogoAdd / LogoParam to control explicit watermarks. It does not use Watermark, WmPosition, or WmUrl.

Request Examples ​

Text-to-video ​

bash
curl -X POST "${BASE_URL}/v1/videos" \
  -H "Authorization: Bearer ${API_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "viduq3-turbo",
    "prompt": "Santa Claus and a bear hug beside a lake, snow falls slowly, slight camera push-in, warm cinematic lighting",
    "duration": 5,
    "metadata": {
      "AspectRatio": "16:9",
      "Audio": true
    }
  }'

First-frame Image-to-video ​

bash
curl -X POST "${BASE_URL}/v1/videos" \
  -H "Authorization: Bearer ${API_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "viduq3-pro",
    "prompt": "The person naturally raises their head and looks at the camera, background lights flow subtly, keep the face clear",
    "image": "https://cdn.example.com/first-frame.png",
    "duration": 5,
    "metadata": {
      "Audio": true,
      "AudioType": "all"
    }
  }'

Start/end-frame Image-to-video ​

When 2 images are provided, the first image is used as the start frame and the second as the end frame. The resolutions of the two images should be close; the recommended start-frame resolution / end-frame resolution ratio is between 0.8 and 1.25.

json
{
  "model": "viduq3-turbo",
  "prompt": "Transition naturally from the start frame to the end frame with continuous motion and smooth camera movement",
  "images": [
    "https://cdn.example.com/start.png",
    "https://cdn.example.com/end.png"
  ],
  "duration": 5,
  "metadata": {
    "Audio": true
  }
}

Reference-to-video ​

Passing 3 to 7 images automatically uses SubmitReferenceToVideoViduJob. You can also use metadata.Action=referenceGenerate to force reference-to-video.

bash
curl -X POST "${BASE_URL}/v1/videos" \
  -H "Authorization: Bearer ${API_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "viduq2",
    "prompt": "Preserve the appearance of the person in the reference images while they walk naturally on a futuristic city street",
    "images": [
      "https://cdn.example.com/character-front.png",
      "https://cdn.example.com/character-side.png",
      "https://cdn.example.com/scene.png"
    ],
    "duration": 5,
    "metadata": {
      "AspectRatio": "16:9"
    }
  }'

Reference Subjects ​

Subjects is used for subject calls. It supports 1 to 7 subjects, with 1 to 7 total subject images; each subject can have at most 3 images. You can reference subjects in the prompt with @subject_id.

json
{
  "model": "viduq2",
  "prompt": "@subject_1 and @subject_2 talk at a street-side cafe, with a voiceover saying the weather is nice today",
  "duration": 5,
  "metadata": {
    "Action": "referenceGenerate",
    "Subjects": [
      {
        "Id": "subject_1",
        "Name": "subject_1",
        "Images": [
          "https://cdn.example.com/person-a-front.png",
          "https://cdn.example.com/person-a-side.png"
        ],
        "VoiceId": "male-qn-qingse"
      },
      {
        "Id": "subject_2",
        "Name": "subject_2",
        "Images": [
          "https://cdn.example.com/person-b-front.png"
        ]
      }
    ],
    "Audio": true,
    "AudioType": "speech_only"
  }
}

Reference subject fields:

FieldTypeRequiredDescription
IdstringYesSubject ID, referenced in the prompt as @Id
Imagesstring[]YesSubject image URLs; each subject can have up to 3 images
NamestringNoSubject name, usually the same as Id
Videosstring[]NoSubject video URLs; supported only by viduq2-pro
VoiceIdstringNoSubject voice ID

Video Reference ​

Video reference is passed through via metadata.Videos. This capability is supported only by viduq2-pro, with up to one 8-second video or two 5-second videos. Supported formats are mp4, avi, and mov; size must not exceed 100 MB.

json
{
  "model": "viduq2",
  "prompt": "Use the character's motion rhythm from the reference video to generate a subject-consistent video",
  "duration": 5,
  "metadata": {
    "Action": "referenceGenerate",
    "Videos": [
      "https://cdn.example.com/reference-motion.mp4"
    ]
  }
}

Callback, Off-peak Mode, Watermark, and Business Payload ​

json
{
  "model": "viduq3-pro",
  "prompt": "A product commercial shot with the subject rotating slowly against soft background lighting",
  "image": "https://cdn.example.com/product.png",
  "duration": 5,
  "metadata": {
    "Payload": "client-order-20260909-0001",
    "CallbackUrl": "https://example.com/callbacks/video",
    "OffPeak": true,
    "LogoAdd": 0
  }
}

The upstream callback does not replace task querying through this service. Clients should still poll status using the task_id returned by this service.

Input Asset Constraints ​

AssetConstraints
Image-to-video imageSupports URL or Base64; formats png, jpeg, jpg, webp; size no more than 50 MB; avoid aspect ratios over 1:4 or 4:1
Start/end-frame imagesPass 2 images: the first is the start frame and the second is the end frame. The two image resolutions must be close, with start-frame resolution / end-frame resolution between 0.8 and 1.25
Reference imagesimages supports 1 to 7 images; formats png, jpeg, jpg, webp; pixels at least 128x128; size no more than 50 MB
Reference subject imagesSubjects supports 1 to 7 subjects, with 1 to 7 total subject images; each subject can have at most 3 images
Reference videosSupports one 8-second video or two 5-second videos; formats mp4, avi, mov; pixels at least 128x128; size no more than 100 MB

Watermark Parameters ​

json
{
  "metadata": {
    "LogoAdd": 1,
    "LogoParam": {
      "LogoUrl": "https://cdn.example.com/logo.png",
      "LogoRect": {
        "X": -222,
        "Y": -54,
        "Width": 202,
        "Height": 34
      }
    }
  }
}
FieldTypeDescription
LogoUrlstringWatermark image URL
LogoImagestringWatermark image Base64; if both LogoImage and LogoUrl are provided, LogoUrl takes precedence
LogoRect.XintegerWatermark rectangle X coordinate; positive values move from left to right, negative values from right to left
LogoRect.YintegerWatermark rectangle Y coordinate; positive values move from top to bottom, negative values from bottom to top
LogoRect.WidthintegerWatermark rectangle width in px
LogoRect.HeightintegerWatermark rectangle height in px

Submission Response ​

After POST /v1/videos succeeds, the service returns a public task ID. The task runs asynchronously and must be queried later.

json
{
  "id": "task_7a31e02c4b",
  "task_id": "task_7a31e02c4b",
  "object": "video",
  "model": "viduq3-pro",
  "status": "queued",
  "progress": 0,
  "created_at": 1788912000
}

The upstream JobId in the original submission response is mapped by this service to the upstream ID of the internal task. Clients only need to save the task_id returned by this service.

Query a Task ​

bash
curl "${BASE_URL}/v1/videos/task_7a31e02c4b" \
  -H "Authorization: Bearer ${API_TOKEN}"

Example response when completed:

json
{
  "id": "task_7a31e02c4b",
  "object": "video",
  "model": "viduq3-pro",
  "status": "completed",
  "progress": 100,
  "created_at": 1788912000,
  "completed_at": 1788912060,
  "metadata": {
    "url": "https://example.com/generated-video.mp4"
  }
}

Upstream query statuses are mapped to this service's task statuses:

Upstream StatusMeaningService status
WAITWaitingqueued / submitted
RUNRunningin_progress
DONETask succeededcompleted
FAILTask failedfailed

The upstream ResultVideoUrl is valid for 24 hours. If COS persistence is configured, this service copies the upstream temporary URL to its own COS URL. If COS is not configured or persistence fails, the upstream temporary URL is returned.

Polling Example ​

bash
TASK_ID="task_7a31e02c4b"

while true; do
  BODY=$(curl -s "${BASE_URL}/v1/videos/${TASK_ID}" \
    -H "Authorization: Bearer ${API_TOKEN}")
  echo "${BODY}"

  STATUS=$(printf '%s' "${BODY}" | jq -r '.status')
  if [ "${STATUS}" = "completed" ] || [ "${STATUS}" = "failed" ]; then
    break
  fi

  sleep 5
done

Poll every 3 to 5 seconds. Do not create duplicate video tasks for the same business request.

Download the Video ​

bash
curl -L "${BASE_URL}/v1/videos/task_7a31e02c4b/content" \
  -H "Authorization: Bearer ${API_TOKEN}" \
  -o output.mp4

If metadata.url in the query response or result_url in the legacy response is a temporary URL, it may expire. For long-term retention, download the video and copy it to your own object storage.

Common Errors ​

SymptomPossible causeSuggested handling
unsupported modelThe model name is not in this service's model list, or the administrator has not enabled the modelCheck the model spelling and backend model configuration
Text-to-video request rejectedAn image-to-video-only model was used, such as viduq2-pro or viduq2-turboSwitch to viduq2, viduq3-pro, or viduq3-turbo
Image-to-video request rejectedA text-to-video-only model was used, or the image URL / Base64 is invalidSwitch to an image-to-video model; check image format, size, aspect ratio, and accessibility
Base64 image failedVidu Actions require URLs, COS conversion is not configured in this service, or the Base64 cannot be parsedUse a public HTTPS image directly, or ask an administrator to configure COS
Reference-to-video failedA non-viduq2 model was used, image/subject/video count is invalid, or asset format is invalidPrefer viduq2; reduce input count or replace assets according to the asset constraints
Resolution does not take effectThis service currently filters metadata.Resolution / metadata.resolutionUse the default resolution for now; adapter changes are required to support it
Watermark does not take effectVidu uses LogoAdd / LogoParamUse metadata.LogoAdd=0 or pass LogoParam
Audio not generatedThe model or capability does not support Audio, or AudioType / VoiceId does not matchPrefer Q3 series for text-to-video; enable audio for reference-to-video only in subject calls
Query never completesThe upstream task is still in WAIT / RUN, or OffPeak placed it in an off-peak queueKeep polling; off-peak mode may take up to 48 hours
Download failedUpstream URL expired or was blocked by proxy security policyDownload promptly; ask an administrator to check COS persistence and the video proxy if needed