Skip to content
Scalekit Docs

Connect AI agents to the Kling AI MCP server

Vendor MCP10 toolsOAuth 2.1/DCRAIMediaDesign

The Kling AI MCP connector routes your AI agent's tool calls to Kling AI's own MCP server through Scalekit. Each user signs in to Kling AI once, and Scalekit stores and refreshes their tokens, so your agent never handles credentials. It comes with 10 tools.

  1. Terminal window
    npm install @scalekit-sdk/node dotenv

    Full SDK reference: Node.js | Python

  2. Add your Scalekit credentials to your .env file. Find values in app.scalekit.com > Developers > API Credentials.

    .env
    SCALEKIT_ENVIRONMENT_URL=<your-environment-url>
    SCALEKIT_CLIENT_ID=<your-client-id>
    SCALEKIT_CLIENT_SECRET=<your-client-secret>
  3. quickstart.mts
    import { ScalekitClient } from '@scalekit-sdk/node'
    import 'dotenv/config'
    import { createInterface } from 'node:readline/promises'
    const scalekit = new ScalekitClient(
    process.env.SCALEKIT_ENVIRONMENT_URL,
    process.env.SCALEKIT_CLIENT_ID,
    process.env.SCALEKIT_CLIENT_SECRET,
    )
    const actions = scalekit.actions
    const connector = 'klingmcp'
    const identifier = 'user_123'
    // Generate an authorization link for the user
    const { link } = await actions.getAuthorizationLink({ connectionName: connector, identifier })
    console.log('Authorize Kling AI MCP:', link)
    const rl = createInterface({ input: process.stdin, output: process.stdout })
    await rl.question('Press Enter after authorizing...')
    rl.close()
    // Make your first call
    const result = await actions.executeTool({
    connector,
    identifier,
    toolName: 'klingmcp_kling_list_actions',
    toolInput: {},
    })
    console.log(result)
    Terminal window
    npx tsx quickstart.mts

Connect this agent connector to let your agent:

  • Photo kling talking — Animate a portrait photo to match a provided audio track (talking-photo)
  • Sync kling lip — Synchronize lip movements in a video to match a given audio track or text
  • List kling — List all available Kling models for video generation
  • Get kling — Query multiple video generation tasks at once
  • Image kling generate video from — Generate AI video using reference images as start and/or end frames
  • Video kling generate, kling extend — Generate AI video from a text prompt using Kling

Use the exact tool names from the Tool list below when you call execute_tool. If you’re not sure which name to use, list the tools available for the current user first.

klingmcp_kling_extend_video#Extend an existing video with additional content.7 params

Extend an existing video with additional content.

NameTypeRequiredDescription
promptstringrequiredDescription of what should happen in the extended portion of the video.
video_idstringrequiredID of the video to extend.
cfg_scalenumberoptionalClassifier-free guidance scale controlling prompt adherence.
durationintegeroptionalDuration of the extended segment in seconds.
modestringoptionalGeneration mode controlling quality and resolution.
modelstringoptionalKling model to use for the extension.
negative_promptstringoptionalThings to avoid in the extended video.
klingmcp_kling_generate_motion#Transfer motion from a reference video to a character image.9 params

Transfer motion from a reference video to a character image.

NameTypeRequiredDescription
image_urlstringrequiredURL of the character image to animate.
video_urlstringrequiredURL of the reference video providing the motion to transfer.
callback_urlstringoptionalWebhook URL to receive task completion notifications.
character_orientationstringoptionalOrientation source for the character.
keep_original_soundstringoptionalWhether to keep the original sound from the reference video ('yes' or 'no').
modestringoptionalGeneration mode controlling quality and resolution.
model_namestringoptionalOptional Kling motion model name, such as 'kling-v2-6' or 'kling-v3'.
promptstringoptionalOptional text description to guide the motion transfer.
watermark_infoobjectoptionalOptional watermark configuration forwarded to the Kling API.
klingmcp_kling_generate_video#Generate AI video from a text prompt using Kling.12 params

Generate AI video from a text prompt using Kling.

NameTypeRequiredDescription
promptstringrequiredDescription of the video to generate.
aspect_ratiostringoptionalVideo aspect ratio.
callback_urlstringoptionalWebhook URL to receive task completion notifications.
camera_controlstringoptionalCamera control parameters as a JSON string.
cfg_scalenumberoptionalClassifier-free guidance scale controlling prompt adherence.
durationintegeroptionalVideo duration in seconds.
generate_audiobooleanoptionalWhether to generate audio synchronously alongside the video.
image_listarrayoptionalAdditional Omni reference images used by kling-v3-omni or kling-o1 models.
modestringoptionalGeneration mode controlling quality and resolution.
modelstringoptionalKling model to use for video generation.
negative_promptstringoptionalThings to avoid in the generated video.
video_listarrayoptionalOne Omni reference video used by kling-v3-omni or kling-o1 models.
klingmcp_kling_generate_video_from_image#Generate AI video using reference images as start and/or end frames.14 params

Generate AI video using reference images as start and/or end frames.

NameTypeRequiredDescription
promptstringrequiredDescription of the video motion and content.
aspect_ratiostringoptionalVideo aspect ratio.
callback_urlstringoptionalWebhook URL to receive task completion notifications.
camera_controlstringoptionalCamera control parameters as a JSON string.
cfg_scalenumberoptionalClassifier-free guidance scale controlling prompt adherence.
durationintegeroptionalVideo duration in seconds.
end_image_urlstringoptionalURL of the image to use as the last frame of the video.
generate_audiobooleanoptionalWhether to generate audio synchronously alongside the video.
image_listarrayoptionalAdditional Omni reference images used by kling-v3-omni or kling-o1 models.
modestringoptionalGeneration mode controlling quality and resolution.
modelstringoptionalKling model to use for video generation.
negative_promptstringoptionalThings to avoid in the generated video.
start_image_urlstringoptionalURL of the image to use as the first frame of the video.
video_listarrayoptionalOne Omni reference video used by kling-v3-omni or kling-o1 models.
klingmcp_kling_get_task#Query the status and result of a video generation task.1 param

Query the status and result of a video generation task.

NameTypeRequiredDescription
task_idstringrequiredThe task ID returned from a generation request.
klingmcp_kling_get_tasks_batch#Query multiple video generation tasks at once.1 param

Query multiple video generation tasks at once.

NameTypeRequiredDescription
task_idsarrayrequiredList of task IDs to query. Maximum recommended batch size is 50 tasks.
klingmcp_kling_lip_sync#Synchronize lip movements in a video to match a given audio track or text.11 params

Synchronize lip movements in a video to match a given audio track or text.

NameTypeRequiredDescription
modestringrequiredLip-sync mode: 'audio2video' to drive lips from an audio file/URL, or 'text2video' to generate speech from text and drive the video.
audio_filestringoptionalBase64-encoded audio file content. Required when mode='audio2video' and audio_type='file'.
audio_typestringoptionalAudio source type. 'url' (default) to supply audio_url; 'file' to supply audio_file as a base64-encoded string.
audio_urlstringoptionalURL of the driving audio. Required when mode='audio2video' and audio_type='url'.
callback_urlstringoptionalWebhook URL that receives a POST when the lip-sync task completes.
textstringoptionalText to convert to speech. Required when mode='text2video'.
video_idstringoptionalTask ID of a previously generated video to use as the source. Provide either video_url or video_id.
video_urlstringoptionalURL of the source video whose lip movements will be replaced. Provide either video_url or video_id.
voice_idstringoptionalVoice ID to use for text-to-speech synthesis (mode='text2video').
voice_languagestringoptionalLanguage of the TTS voice. 'zh' for Chinese (default), 'en' for English. Used when mode='text2video'.
voice_speednumberoptionalSpeech speed multiplier from 0.8 to 2.0 (default 1.0). Used when mode='text2video'.
klingmcp_kling_list_actions#List all available Kling API actions and corresponding tools.0 params

List all available Kling API actions and corresponding tools.

klingmcp_kling_list_models#List all available Kling models for video generation.0 params

List all available Kling models for video generation.

klingmcp_kling_talking_photo#Animate a portrait photo to match a provided audio track (talking-photo).7 params

Animate a portrait photo to match a provided audio track (talking-photo).

NameTypeRequiredDescription
audio_urlstringrequiredURL of the audio file that drives the talking animation.
image_urlstringrequiredURL of the portrait image to animate. Should be a clear frontal face photo.
callback_urlstringoptionalWebhook URL that receives a POST when the talking-photo task completes.
durationintegeroptionalVideo duration in seconds. Options: 5 (default) or 10.
modestringoptionalGeneration quality mode. 'pro' (default) for higher quality; 'std' for faster generation.
modelstringoptionalKling model version. Default is 'kling-v2-1-master'.
promptstringoptionalOptional text description to guide the animation style or content.