Skip to content

Kling Video Generation API ​

This guide explains how to call Kling video generation models through this service. Use an API token issued by this service, and replace BASE_URL and API_TOKEN in the examples with your actual values.

Put all extension parameters in metadata and use PascalCase field names, such as Sound, ImageTail, and CameraControl. Do not pass Model, Prompt, Image, Duration, or Mode in metadata; these fields are generated from top-level request parameters.

Endpoint Overview ​

MethodPathPurposeResponse style
POST/v1/videosCreate a video generation taskOpenAI Video object
GET/v1/videos/{task_id}Query task status and resultOpenAI Video object
GET/v1/videos/{task_id}/contentDownload the generated videoVideo binary
POST/v1/video/generationsCreate a video generation taskLegacy-compatible format
GET/v1/video/generations/{task_id}Query task status and resultLegacy-compatible format
POST/v1/videos/generationsCreate a video generation taskLegacy-compatible format

Capability Routing ​

This service automatically selects text-to-video or image-to-video based on image input:

Request shapeCapabilityNotes
No image / imagesText-to-videoGenerate video from text prompt only
Includes imageImage-to-videoUses image as the first frame or reference image
Includes imagesImage-to-videoOnly the first image is used as the first frame or reference image; use metadata.ImageTail for an end frame

Supported Models ​

API model nameUpstream model codeText-to-videoImage-to-video
kling-v1v1.0SupportedNot recommended
kling-v1-5v1.5SupportedNot recommended
kling-v1-6v1.6SupportedSupported
kling-v2-masterv2.0SupportedSupported
kling-v2-1v2.1Not recommendedSupported
kling-v2-1-masterv2.1mSupportedNot recommended
kling-v2-5-turbov2.5SupportedSupported
kling-v2-6v2.6SupportedSupported
kling-v3v3.0SupportedSupported

kling-v2-1 corresponds to v2.1 in the image-to-video documentation; kling-v2-1-master corresponds to v2.1m in the text-to-video documentation. For new integrations, choose the model by capability and do not use kling-v2-1-master for image-to-video.

Authentication ​

All requests use a Bearer token:

http
Authorization: Bearer sk-...
Content-Type: application/json

Request Parameters ​

Top-level Parameters ​

ParameterTypeRequiredDescription
modelstringYesKling model name to call. It must be an API model name from the "Supported Models" table, such as kling-v2-6 or kling-v3. The service converts it to the upstream model code.
promptstringYesVideo content description. Describe the subject, action, scene, camera language, and style when possible, for example: "A man walks past neon signs with an umbrella on a rainy night, slow tracking shot, cinematic." Recommended maximum length: 2500 characters.
imagestringRequired for image-to-videoSingle first frame or reference image. Supports public HTTP/HTTPS URLs or Base64 images. When provided, the request is handled as image-to-video. Recommended image size is no more than 10 MB, resolution at least 300x300, aspect ratio between 1:2.5 and 2.5:1, and format JPG, JPEG, or PNG.
imagesstring[]NoArray of image URLs/Base64 images. Current Kling image-to-video uses only the first image. For an end frame, do not put the second image in images; use metadata.ImageTail instead.
durationinteger/stringNoVideo duration in seconds. Defaults to 5 when omitted. See the Duration section for available values.
modestringNoGeneration mode. Defaults to std when omitted. See the Mode section for available values.
metadataobjectNoExtension parameter object. Only put fields from the "metadata Parameters" table below, using PascalCase field names.

metadata Parameters ​

Do not include Model, Prompt, Image, Duration, or Mode in metadata; these fields are generated from top-level model, prompt, image / images, duration, and mode. Duplicating them can cause mismatched request content and billing model.

ParameterTypeRequiredApplicable models / capabilitiesDescription
ImageTailobjectNokling-v1-6, kling-v2-master, kling-v2-1, kling-v2-5-turbo, kling-v2-6, kling-v3; image-to-video onlyEnd frame image. Object fields: Url string, public image URL; Base64 string, image Base64. Provide one of them. Image requirements are the same as image. Do not use together with CameraControl, StaticMask, or DynamicMasks.
AspectRatiostringNoAll text-to-video modelsOutput aspect ratio. Available values: 16:9, 9:16, 1:1; defaults to 16:9 when omitted. Image-to-video uses the input image aspect ratio, so do not pass this field.
NegativePromptstringNoAll models; text-to-video and image-to-videoNegative prompt, used to reduce blur, distortion, low clarity, and similar issues. Recommended maximum length: 2500 characters.
CfgScalenumberNokling-v1, kling-v1-5, kling-v1-6, kling-v2-1, kling-v2-1-master, kling-v3Prompt adherence strength, range [0,1]; defaults to 0.5 when omitted. Do not use with kling-v2-master, kling-v2-5-turbo, or kling-v2-6.
SoundstringNokling-v2-6, kling-v3; text-to-video and image-to-videoWhether to generate audio. Available values: on, off. kling-v2-6 is silent when using mode=std; use mode=pro or omit mode when audio is needed.
VoiceListarray<object>Nokling-v2-6; image-to-video onlyList of voices. Requires Sound=on. Up to 2 objects, each containing VoiceId string. The corresponding voice ID must be referenced in Prompt. Do not use with ElementList; kling-v3 does not support specifying voices.
CameraControlobjectNokling-v1-6, kling-v2-master, kling-v2-1, kling-v2-1-master, kling-v2-5-turbo, kling-v2-6, kling-v3; text-to-video and image-to-videoCamera control. Object fields: Type string, available values simple, down_back, forward_up, right_turn_forward, left_turn_forward; Config object. When Type=simple, Config is required and exactly one of Horizontal, Vertical, Pan, Tilt, Roll, or Zoom should be provided. All are numbers in range [-10,10]. For image-to-video, do not use together with ImageTail, StaticMask, or DynamicMasks.
StaticMaskstringNokling-v1-6, kling-v2-master, kling-v2-1, kling-v2-5-turbo, kling-v2-6, kling-v3; image-to-video onlyStatic mask image, as a public URL or Base64. Format requirements are the same as image; aspect ratio must match image. If used together with DynamicMasks, resolution must match DynamicMasks.Mask. Do not use together with ImageTail or CameraControl.
DynamicMasksarray<object>Nokling-v1-6, kling-v2-master, kling-v2-1, kling-v2-5-turbo, kling-v2-6, kling-v3; image-to-video onlyDynamic mask list, up to 6 objects. Each object contains: Mask string, mask image URL/Base64 with the same format requirements as image and the same aspect ratio as image; Trajectories array<object>, motion trajectory points. Each trajectory point contains X integer and Y integer. For a 5-second video, the trajectory point count is 2 to 77. The coordinate origin is the lower-left corner of the image. Do not use together with ImageTail or CameraControl.
MultiShotbooleanNokling-v3; text-to-video and image-to-videoWhether to enable multi-shot generation. When set to true, Prompt does not take effect and ShotType must also be provided. For image-to-video, do not use together with ImageTail.
ShotTypestringNokling-v3; text-to-video and image-to-videoShot splitting mode. Available values: customize, intelligence. When using customize, MultiPrompt is required.
MultiPromptarray<object>Nokling-v3; text-to-video and image-to-videoCustom shot prompts, 1 to 6 objects. Each object contains Index integer, Prompt string, and Duration string. Each shot prompt can be up to 512 characters. Shot duration must be at least 1 second and no longer than the total task duration. The sum of all shot Duration values must equal the total task duration.
ElementListarray<object>Nokling-v3; image-to-video onlyReference subject list, up to 3 objects. Each object contains ElementId string, representing the subject ID in the subject library. Do not use with VoiceList.
LogoAddboolean/integerNoAll models; text-to-video and image-to-videoWhether to add a watermark or AI label. Pass false or 0 to disable it.
LogoParamobjectNoAll models; text-to-video and image-to-videoWatermark parameter object. Fields: LogoUrl string, watermark image URL; LogoImage string, watermark image Base64. Provide one of LogoUrl or LogoImage; if both are provided, LogoUrl takes precedence. LogoRect object contains X, Y, Width, and Height integer fields in px.
CallbackUrlstringNoAll models; text-to-video and image-to-videoUpstream callback URL. The upstream callback does not replace task querying through this service.
ExternalTaskIdstringNoAll models; text-to-video and image-to-videoExternal task ID for business idempotency or tracing.

Duration ​

API model nameText-to-video valuesImage-to-video values
kling-v15, 10Not recommended
kling-v1-5Recommended to omit and use the upstream defaultNot recommended
kling-v1-65, 105, 10
kling-v2-master5, 105, 10
kling-v2-1Not recommended5, 10
kling-v2-1-master5, 10Not recommended
kling-v2-5-turbo5, 105, 10
kling-v2-65, 105, 10
kling-v3Integers from 3 to 15Integers from 3 to 15

Mode ​

API model nameText-to-video valuesImage-to-video values
kling-v1proNot recommended
kling-v1-5proNot recommended
kling-v1-6std, proUse pro for first-frame or start/end-frame requests
kling-v2-masterRecommended to omitRecommended to omit
kling-v2-1Not recommendedUse pro for start/end-frame requests
kling-v2-1-masterRecommended to omitNot recommended
kling-v2-5-turbopro can be used for start/end-frame requests; otherwise omit itUse pro for start/end-frame requests
kling-v2-6pro or omitted; do not use std when audio is neededUse pro for start/end-frame requests; do not use std when audio is needed
kling-v3Recommended to omitRecommended to omit

Request Examples ​

Text-to-video ​

bash
curl -X POST "${BASE_URL}/v1/videos" \
  -H "Authorization: Bearer ${API_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kling-v2-6",
    "prompt": "At sunrise, a white sailboat slowly crosses a bay, cinematic shot, soft golden light",
    "duration": 5,
    "mode": "pro"
  }'

Image-to-video ​

bash
curl -X POST "${BASE_URL}/v1/videos" \
  -H "Authorization: Bearer ${API_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kling-v2-6",
    "prompt": "Slowly push the camera forward while the sea and clouds move naturally, keeping the subject composition stable",
    "image": "https://cdn.example.com/first-frame.png",
    "duration": 5,
    "mode": "pro"
  }'

Direct Audio/Video Output ​

json
{
  "model": "kling-v3",
  "prompt": "A cinematic rainy-night street shot, a character speaks in a low voiceover, car lights sweep through the background",
  "image": "https://cdn.example.com/start.png",
  "duration": 5,
  "metadata": {
    "Sound": "on"
  }
}

Start and End Frames ​

json
{
  "model": "kling-v2-6",
  "prompt": "Transition naturally from the starting frame to the ending frame with smooth camera motion and continuous character movement",
  "image": "https://cdn.example.com/start.png",
  "duration": 5,
  "mode": "pro",
  "metadata": {
    "ImageTail": {
      "Url": "https://cdn.example.com/end.png"
    }
  }
}

Camera Control ​

json
{
  "model": "kling-v2-6",
  "prompt": "Slowly push in from a wide shot to the subject's face with cinematic depth of field",
  "image": "https://cdn.example.com/portrait.png",
  "duration": 5,
  "mode": "pro",
  "metadata": {
    "NegativePrompt": "blurry, distorted, low clarity, camera shake",
    "CameraControl": {
      "Type": "simple",
      "Config": {
        "Zoom": 2.0,
        "Vertical": 0.2
      }
    }
  }
}

Motion Brush ​

json
{
  "model": "kling-v2-6",
  "prompt": "Only make the character's arm wave slightly while keeping the background stable",
  "image": "https://cdn.example.com/portrait.png",
  "duration": 5,
  "metadata": {
    "DynamicMasks": [
      {
        "Mask": "https://cdn.example.com/arm-mask.png",
        "Trajectories": [
          { "X": 410, "Y": 520 },
          { "X": 460, "Y": 500 }
        ]
      }
    ]
  }
}

Multi-shot and Subject Reference ​

json
{
  "model": "kling-v3",
  "prompt": "A character walks from outdoors into a coffee shop while preserving character consistency",
  "duration": 10,
  "metadata": {
    "MultiShot": true,
    "ShotType": "customize",
    "MultiPrompt": [
      { "Prompt": "Wide shot, the character walks past a street corner" },
      { "Prompt": "Medium shot, the character pushes the door open and enters the coffee shop" }
    ],
    "ElementList": [
      {
        "Name": "main_character",
        "Image": {
          "Url": "https://cdn.example.com/character.png"
        }
      }
    ]
  }
}

Submission Response ​

After POST /v1/videos succeeds, the service returns a public task ID. The task runs asynchronously and must be queried later.

json
{
  "id": "task_2c8f7b3e9a",
  "task_id": "task_2c8f7b3e9a",
  "object": "video",
  "model": "kling-v2-6",
  "status": "queued",
  "progress": 0,
  "created_at": 1788912000
}

Query a Task ​

bash
curl "${BASE_URL}/v1/videos/task_2c8f7b3e9a" \
  -H "Authorization: Bearer ${API_TOKEN}"

Example response when completed:

json
{
  "id": "task_2c8f7b3e9a",
  "object": "video",
  "model": "kling-v2-6",
  "status": "completed",
  "progress": 100,
  "created_at": 1788912000,
  "completed_at": 1788912042,
  "metadata": {
    "url": "https://example.com/generated-video.mp4"
  }
}

The result_url in legacy responses points to the same type of result video as metadata.url in OpenAI Video responses.

Polling Example ​

bash
TASK_ID="task_2c8f7b3e9a"

while true; do
  BODY=$(curl -s "${BASE_URL}/v1/videos/${TASK_ID}" \
    -H "Authorization: Bearer ${API_TOKEN}")
  echo "${BODY}"

  STATUS=$(printf '%s' "${BODY}" | jq -r '.status')
  if [ "${STATUS}" = "completed" ] || [ "${STATUS}" = "failed" ]; then
    break
  fi

  sleep 5
done

Poll every 3 to 5 seconds. Do not create duplicate video tasks for the same business request.

Download the Video ​

bash
curl -L "${BASE_URL}/v1/videos/task_2c8f7b3e9a/content" \
  -H "Authorization: Bearer ${API_TOKEN}" \
  -o output.mp4

If metadata.url or result_url in the query response is a temporary URL, it may expire. For long-term retention, download the video and copy it to your own object storage.

Common Errors ​

SymptomPossible causeSuggested handling
unsupported modelThe model name is not in the supported list, or the channel has not enabled the modelCheck the model spelling and backend model configuration
Image URL rejectedThe image is not publicly accessible, or its format/dimensions do not meet requirementsUse a public HTTPS image and check image size, format, and aspect ratio
Parameter rejected by upstreamFields do not match the "metadata Parameters" table, or mutually exclusive fields are used togetherRemove unsupported fields according to the table; keep only one of end frame, camera control, static mask, and dynamic mask
Audio not generatedThe model is not kling-v2-6 / kling-v3, or kling-v2-6 uses mode=stdSwitch to kling-v2-6 / kling-v3, and use mode=pro or omit the mode
Query never completesUpstream queue congestion or long model generation timeKeep polling the same task ID and avoid duplicate submissions
Download failedResult URL expired or was blocked by proxy security policyDownload promptly; ask an administrator to check the video proxy and SSRF configuration if needed