APIXO maps this route to Kuaishou’s Kling Video 2.6 and does not expose a Standard, Pro, Master, or Turbo quality-tier selector.
Kuaishou’s Kling Video 2.6 generates short video and optional synchronized audio in one pass. Through APIXO, developers can create from text or one starting image, add an optional end frame, select five- or ten-second output, and choose between silent generation or speech, effects, ambience, and music.
APIXO Workflows
Text-to-video / Image-to-video
Silent Price
From $0.30 / video5 seconds without sound
With Sound
From $0.60 / video5 seconds with generated audio
Duration
5 or 10 seconds
Released
Dec 3, 2025
Loading workspace...
Kling Video 2.6 generates visuals and sound within the same model pass. This enables closer coordination between visible action, vocal rhythm, environmental events, and the timing of the resulting soundtrack.
Direct standalone or combined speech, dialogue, narration, singing, rap, music, ambient sound, and effects in the prompt. APIXO’s sound toggle determines whether the requested output includes generated audio.
Kling Video 2.6 supports dialogue-led scenes involving multiple characters, including interviews, scripted exchanges, and short performances. Speaker attribution, pronunciation, and synchronization should still be reviewed before publishing.
The model interprets textual descriptions, colloquial expressions, and short story scenarios to coordinate sound with visual events. Explicitly describe who speaks, what happens, and when each sound should occur.
Combine voice with environmental ambience and event-specific effects instead of producing a single isolated sound. This supports compact scenes where atmosphere, action, and spoken content all contribute to the result.
Through APIXO’s image-to-video workflow, one image establishes the starting composition and an optional second image can guide the intended end frame. The model generates the intervening motion rather than performing deterministic interpolation.
Generate from a required prompt without source images. APIXO exposes 1:1, 9:16, and 16:9 aspect ratios for this workflow, supporting square, vertical, and landscape audiovisual scenes.
Supply one image as the starting frame and optionally a second image as an end frame. APIXO does not forward aspect-ratio selection in this mode, so prepare source images for the intended composition.
The APIXO route exposes optional audio controls without a separate quality-tier selector.
1–1,000 characters
Prompt Length
Up to 500 characters
Negative Prompt
1–2 image URLs
Image Input
JPG / PNG / WebP
Source Formats
0–1
CFG Scale
MP4 video URL
APIXO Result
Generate product demonstrations with narration, environmental sound, and action-specific effects in one pass. Use an image-led workflow when an approved product composition should anchor the shot, then verify labels, geometry, speech, and claims.
Create interviews, short scripted exchanges, comedy concepts, singing, or rap performances. Kling Video 2.6 supports generated dialogue, but speaker attribution, pronunciation, timing, and lip synchronization should be reviewed before publishing.
Prototype how camera movement, physical action, dialogue, ambience, and effects interact within a five- or ten-second scene. This helps teams evaluate the full audiovisual idea before separate production or postproduction.
Animate one approved image or guide the clip toward an optional ending frame. Add sound when the transition needs synchronized atmosphere or effects, or generate silently for an existing audio workflow.
APIXO exposes one Kling 2.6 variant rather than Standard, Pro, Master, or Turbo quality tiers.
Every request requires a prompt, a 5- or 10-second duration, and an explicit sound value.
Text-to-video accepts aspect-ratio selection; image-to-video derives composition from its source image.
Image-to-video accepts one starting image and one optional end-frame image.
The route does not expose video input, motion control, extension, editing, or arbitrary multi-reference generation.
APIXO supports status polling or webhook delivery; Longer clips with generated audio typically require more processing time; callback delivery is preferable for production workloads.
APIXO maps this route to Kuaishou’s Kling Video 2.6 and does not expose a Standard, Pro, Master, or Turbo quality-tier selector.
Silent output costs $0.06 per generated second: $0.30 for five seconds or $0.60 for ten. With model-generated audio enabled, the rate is $0.12 per second: $0.60 or $1.20 respectively.
Yes. Kuaishou officially supports speech, dialogue, narration, singing, and rap alongside effects and ambience. The model generates Chinese and English voices, while APIXO exposes this capability through the required sound toggle and prompt.
No. APIXO does not expose audio-upload or video-input fields for this route. Kling 2.6 generates requested sound from the prompt; it does not accept a prerecorded voice, soundtrack, motion clip, or source video here.
Check identity, anatomy, hands, embedded text, object contact, rapid movement, transitions, dialogue accuracy, speaker assignment, pronunciation, and audio synchronization. Multi-character scenes and dense sound directions may require several generations.
Discover more AI models for your next creative workflow