Kling Video Generation API
This guide explains how to call Kling video generation models through this service. Use an API token issued by this service, and replace BASE_URL and API_TOKEN in the examples with your actual values.
Put all extension parameters in metadata and use PascalCase field names, such as Sound, ImageTail, and CameraControl. Do not pass Model, Prompt, Image, Duration, or Mode in metadata; these fields are generated from top-level request parameters.
Endpoint Overview
| Method | Path | Purpose | Response style |
|---|---|---|---|
POST | /v1/videos | Create a video generation task | OpenAI Video object |
GET | /v1/videos/{task_id} | Query task status and result | OpenAI Video object |
GET | /v1/videos/{task_id}/content | Download the generated video | Video binary |
POST | /v1/video/generations | Create a video generation task | Legacy-compatible format |
GET | /v1/video/generations/{task_id} | Query task status and result | Legacy-compatible format |
POST | /v1/videos/generations | Create a video generation task | Legacy-compatible format |
Capability Routing
This service automatically selects text-to-video or image-to-video based on image input:
| Request shape | Capability | Notes |
|---|---|---|
No image / images | Text-to-video | Generate video from text prompt only |
Includes image | Image-to-video | Uses image as the first frame or reference image |
Includes images | Image-to-video | Only the first image is used as the first frame or reference image; use metadata.ImageTail for an end frame |
Supported Models
| API model name | Upstream model code | Text-to-video | Image-to-video |
|---|---|---|---|
kling-v1 | v1.0 | Supported | Not recommended |
kling-v1-5 | v1.5 | Supported | Not recommended |
kling-v1-6 | v1.6 | Supported | Supported |
kling-v2-master | v2.0 | Supported | Supported |
kling-v2-1 | v2.1 | Not recommended | Supported |
kling-v2-1-master | v2.1m | Supported | Not recommended |
kling-v2-5-turbo | v2.5 | Supported | Supported |
kling-v2-6 | v2.6 | Supported | Supported |
kling-v3 | v3.0 | Supported | Supported |
kling-v2-1 corresponds to v2.1 in the image-to-video documentation; kling-v2-1-master corresponds to v2.1m in the text-to-video documentation. For new integrations, choose the model by capability and do not use kling-v2-1-master for image-to-video.
Authentication
All requests use a Bearer token:
Authorization: Bearer sk-...
Content-Type: application/jsonRequest Parameters
Top-level Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Kling model name to call. It must be an API model name from the "Supported Models" table, such as kling-v2-6 or kling-v3. The service converts it to the upstream model code. |
prompt | string | Yes | Video content description. Describe the subject, action, scene, camera language, and style when possible, for example: "A man walks past neon signs with an umbrella on a rainy night, slow tracking shot, cinematic." Recommended maximum length: 2500 characters. |
image | string | Required for image-to-video | Single first frame or reference image. Supports public HTTP/HTTPS URLs or Base64 images. When provided, the request is handled as image-to-video. Recommended image size is no more than 10 MB, resolution at least 300x300, aspect ratio between 1:2.5 and 2.5:1, and format JPG, JPEG, or PNG. |
images | string[] | No | Array of image URLs/Base64 images. Current Kling image-to-video uses only the first image. For an end frame, do not put the second image in images; use metadata.ImageTail instead. |
duration | integer/string | No | Video duration in seconds. Defaults to 5 when omitted. See the Duration section for available values. |
mode | string | No | Generation mode. Defaults to std when omitted. See the Mode section for available values. |
metadata | object | No | Extension parameter object. Only put fields from the "metadata Parameters" table below, using PascalCase field names. |
metadata Parameters
Do not include Model, Prompt, Image, Duration, or Mode in metadata; these fields are generated from top-level model, prompt, image / images, duration, and mode. Duplicating them can cause mismatched request content and billing model.
| Parameter | Type | Required | Applicable models / capabilities | Description |
|---|---|---|---|---|
ImageTail | object | No | kling-v1-6, kling-v2-master, kling-v2-1, kling-v2-5-turbo, kling-v2-6, kling-v3; image-to-video only | End frame image. Object fields: Url string, public image URL; Base64 string, image Base64. Provide one of them. Image requirements are the same as image. Do not use together with CameraControl, StaticMask, or DynamicMasks. |
AspectRatio | string | No | All text-to-video models | Output aspect ratio. Available values: 16:9, 9:16, 1:1; defaults to 16:9 when omitted. Image-to-video uses the input image aspect ratio, so do not pass this field. |
NegativePrompt | string | No | All models; text-to-video and image-to-video | Negative prompt, used to reduce blur, distortion, low clarity, and similar issues. Recommended maximum length: 2500 characters. |
CfgScale | number | No | kling-v1, kling-v1-5, kling-v1-6, kling-v2-1, kling-v2-1-master, kling-v3 | Prompt adherence strength, range [0,1]; defaults to 0.5 when omitted. Do not use with kling-v2-master, kling-v2-5-turbo, or kling-v2-6. |
Sound | string | No | kling-v2-6, kling-v3; text-to-video and image-to-video | Whether to generate audio. Available values: on, off. kling-v2-6 is silent when using mode=std; use mode=pro or omit mode when audio is needed. |
VoiceList | array<object> | No | kling-v2-6; image-to-video only | List of voices. Requires Sound=on. Up to 2 objects, each containing VoiceId string. The corresponding voice ID must be referenced in Prompt. Do not use with ElementList; kling-v3 does not support specifying voices. |
CameraControl | object | No | kling-v1-6, kling-v2-master, kling-v2-1, kling-v2-1-master, kling-v2-5-turbo, kling-v2-6, kling-v3; text-to-video and image-to-video | Camera control. Object fields: Type string, available values simple, down_back, forward_up, right_turn_forward, left_turn_forward; Config object. When Type=simple, Config is required and exactly one of Horizontal, Vertical, Pan, Tilt, Roll, or Zoom should be provided. All are numbers in range [-10,10]. For image-to-video, do not use together with ImageTail, StaticMask, or DynamicMasks. |
StaticMask | string | No | kling-v1-6, kling-v2-master, kling-v2-1, kling-v2-5-turbo, kling-v2-6, kling-v3; image-to-video only | Static mask image, as a public URL or Base64. Format requirements are the same as image; aspect ratio must match image. If used together with DynamicMasks, resolution must match DynamicMasks.Mask. Do not use together with ImageTail or CameraControl. |
DynamicMasks | array<object> | No | kling-v1-6, kling-v2-master, kling-v2-1, kling-v2-5-turbo, kling-v2-6, kling-v3; image-to-video only | Dynamic mask list, up to 6 objects. Each object contains: Mask string, mask image URL/Base64 with the same format requirements as image and the same aspect ratio as image; Trajectories array<object>, motion trajectory points. Each trajectory point contains X integer and Y integer. For a 5-second video, the trajectory point count is 2 to 77. The coordinate origin is the lower-left corner of the image. Do not use together with ImageTail or CameraControl. |
MultiShot | boolean | No | kling-v3; text-to-video and image-to-video | Whether to enable multi-shot generation. When set to true, Prompt does not take effect and ShotType must also be provided. For image-to-video, do not use together with ImageTail. |
ShotType | string | No | kling-v3; text-to-video and image-to-video | Shot splitting mode. Available values: customize, intelligence. When using customize, MultiPrompt is required. |
MultiPrompt | array<object> | No | kling-v3; text-to-video and image-to-video | Custom shot prompts, 1 to 6 objects. Each object contains Index integer, Prompt string, and Duration string. Each shot prompt can be up to 512 characters. Shot duration must be at least 1 second and no longer than the total task duration. The sum of all shot Duration values must equal the total task duration. |
ElementList | array<object> | No | kling-v3; image-to-video only | Reference subject list, up to 3 objects. Each object contains ElementId string, representing the subject ID in the subject library. Do not use with VoiceList. |
LogoAdd | boolean/integer | No | All models; text-to-video and image-to-video | Whether to add a watermark or AI label. Pass false or 0 to disable it. |
LogoParam | object | No | All models; text-to-video and image-to-video | Watermark parameter object. Fields: LogoUrl string, watermark image URL; LogoImage string, watermark image Base64. Provide one of LogoUrl or LogoImage; if both are provided, LogoUrl takes precedence. LogoRect object contains X, Y, Width, and Height integer fields in px. |
CallbackUrl | string | No | All models; text-to-video and image-to-video | Upstream callback URL. The upstream callback does not replace task querying through this service. |
ExternalTaskId | string | No | All models; text-to-video and image-to-video | External task ID for business idempotency or tracing. |
Duration
| API model name | Text-to-video values | Image-to-video values |
|---|---|---|
kling-v1 | 5, 10 | Not recommended |
kling-v1-5 | Recommended to omit and use the upstream default | Not recommended |
kling-v1-6 | 5, 10 | 5, 10 |
kling-v2-master | 5, 10 | 5, 10 |
kling-v2-1 | Not recommended | 5, 10 |
kling-v2-1-master | 5, 10 | Not recommended |
kling-v2-5-turbo | 5, 10 | 5, 10 |
kling-v2-6 | 5, 10 | 5, 10 |
kling-v3 | Integers from 3 to 15 | Integers from 3 to 15 |
Mode
| API model name | Text-to-video values | Image-to-video values |
|---|---|---|
kling-v1 | pro | Not recommended |
kling-v1-5 | pro | Not recommended |
kling-v1-6 | std, pro | Use pro for first-frame or start/end-frame requests |
kling-v2-master | Recommended to omit | Recommended to omit |
kling-v2-1 | Not recommended | Use pro for start/end-frame requests |
kling-v2-1-master | Recommended to omit | Not recommended |
kling-v2-5-turbo | pro can be used for start/end-frame requests; otherwise omit it | Use pro for start/end-frame requests |
kling-v2-6 | pro or omitted; do not use std when audio is needed | Use pro for start/end-frame requests; do not use std when audio is needed |
kling-v3 | Recommended to omit | Recommended to omit |
Request Examples
Text-to-video
curl -X POST "${BASE_URL}/v1/videos" \
-H "Authorization: Bearer ${API_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"model": "kling-v2-6",
"prompt": "At sunrise, a white sailboat slowly crosses a bay, cinematic shot, soft golden light",
"duration": 5,
"mode": "pro"
}'Image-to-video
curl -X POST "${BASE_URL}/v1/videos" \
-H "Authorization: Bearer ${API_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"model": "kling-v2-6",
"prompt": "Slowly push the camera forward while the sea and clouds move naturally, keeping the subject composition stable",
"image": "https://cdn.example.com/first-frame.png",
"duration": 5,
"mode": "pro"
}'Direct Audio/Video Output
{
"model": "kling-v3",
"prompt": "A cinematic rainy-night street shot, a character speaks in a low voiceover, car lights sweep through the background",
"image": "https://cdn.example.com/start.png",
"duration": 5,
"metadata": {
"Sound": "on"
}
}Start and End Frames
{
"model": "kling-v2-6",
"prompt": "Transition naturally from the starting frame to the ending frame with smooth camera motion and continuous character movement",
"image": "https://cdn.example.com/start.png",
"duration": 5,
"mode": "pro",
"metadata": {
"ImageTail": {
"Url": "https://cdn.example.com/end.png"
}
}
}Camera Control
{
"model": "kling-v2-6",
"prompt": "Slowly push in from a wide shot to the subject's face with cinematic depth of field",
"image": "https://cdn.example.com/portrait.png",
"duration": 5,
"mode": "pro",
"metadata": {
"NegativePrompt": "blurry, distorted, low clarity, camera shake",
"CameraControl": {
"Type": "simple",
"Config": {
"Zoom": 2.0,
"Vertical": 0.2
}
}
}
}Motion Brush
{
"model": "kling-v2-6",
"prompt": "Only make the character's arm wave slightly while keeping the background stable",
"image": "https://cdn.example.com/portrait.png",
"duration": 5,
"metadata": {
"DynamicMasks": [
{
"Mask": "https://cdn.example.com/arm-mask.png",
"Trajectories": [
{ "X": 410, "Y": 520 },
{ "X": 460, "Y": 500 }
]
}
]
}
}Multi-shot and Subject Reference
{
"model": "kling-v3",
"prompt": "A character walks from outdoors into a coffee shop while preserving character consistency",
"duration": 10,
"metadata": {
"MultiShot": true,
"ShotType": "customize",
"MultiPrompt": [
{ "Prompt": "Wide shot, the character walks past a street corner" },
{ "Prompt": "Medium shot, the character pushes the door open and enters the coffee shop" }
],
"ElementList": [
{
"Name": "main_character",
"Image": {
"Url": "https://cdn.example.com/character.png"
}
}
]
}
}Submission Response
After POST /v1/videos succeeds, the service returns a public task ID. The task runs asynchronously and must be queried later.
{
"id": "task_2c8f7b3e9a",
"task_id": "task_2c8f7b3e9a",
"object": "video",
"model": "kling-v2-6",
"status": "queued",
"progress": 0,
"created_at": 1788912000
}Query a Task
curl "${BASE_URL}/v1/videos/task_2c8f7b3e9a" \
-H "Authorization: Bearer ${API_TOKEN}"Example response when completed:
{
"id": "task_2c8f7b3e9a",
"object": "video",
"model": "kling-v2-6",
"status": "completed",
"progress": 100,
"created_at": 1788912000,
"completed_at": 1788912042,
"metadata": {
"url": "https://example.com/generated-video.mp4"
}
}The result_url in legacy responses points to the same type of result video as metadata.url in OpenAI Video responses.
Polling Example
TASK_ID="task_2c8f7b3e9a"
while true; do
BODY=$(curl -s "${BASE_URL}/v1/videos/${TASK_ID}" \
-H "Authorization: Bearer ${API_TOKEN}")
echo "${BODY}"
STATUS=$(printf '%s' "${BODY}" | jq -r '.status')
if [ "${STATUS}" = "completed" ] || [ "${STATUS}" = "failed" ]; then
break
fi
sleep 5
donePoll every 3 to 5 seconds. Do not create duplicate video tasks for the same business request.
Download the Video
curl -L "${BASE_URL}/v1/videos/task_2c8f7b3e9a/content" \
-H "Authorization: Bearer ${API_TOKEN}" \
-o output.mp4If metadata.url or result_url in the query response is a temporary URL, it may expire. For long-term retention, download the video and copy it to your own object storage.
Common Errors
| Symptom | Possible cause | Suggested handling |
|---|---|---|
unsupported model | The model name is not in the supported list, or the channel has not enabled the model | Check the model spelling and backend model configuration |
| Image URL rejected | The image is not publicly accessible, or its format/dimensions do not meet requirements | Use a public HTTPS image and check image size, format, and aspect ratio |
| Parameter rejected by upstream | Fields do not match the "metadata Parameters" table, or mutually exclusive fields are used together | Remove unsupported fields according to the table; keep only one of end frame, camera control, static mask, and dynamic mask |
| Audio not generated | The model is not kling-v2-6 / kling-v3, or kling-v2-6 uses mode=std | Switch to kling-v2-6 / kling-v3, and use mode=pro or omit the mode |
| Query never completes | Upstream queue congestion or long model generation time | Keep polling the same task ID and avoid duplicate submissions |
| Download failed | Result URL expired or was blocked by proxy security policy | Download promptly; ask an administrator to check the video proxy and SSRF configuration if needed |